Welcome back. In the previous lesson, you mapped the stakeholders who control funding, data access, workflow adoption, and AI governance for the authorization-aware cited evidence assistant. That map matters here: Finance can validate cost assumptions, the Strategic Operations product owner can validate how work is actually done, and data and AI-governance owners can establish the non-negotiable conditions under which value is acceptable.
This lesson turns the chosen use case into an initial, testable investment claim. You will estimate its baseline value hypothesis: a transparent statement of the current workflow, the improvement the platform is expected to create, the assumptions behind that expectation, and how the resulting capacity or financial value will be measured. This is not a promise of ROI. It is the starting model that the pilot must either validate, refine, or disprove.
A value hypothesis is a testable business claim
For this capstone, the primary user is a Strategic Operations analyst who must locate approved enterprise material, synthesize it into a decision-ready brief, and let a director verify the underlying evidence. The platform’s purpose is not to make portfolio or funding decisions. It is to reduce the effort of finding and assembling authorized evidence while improving traceability through citations.
A useful value hypothesis connects five elements:
-
A defined unit of work
What repeatable task will improve? Here, it is an eligible decision-ready evidence pack. -
A measured baseline
How long does that task take today, at what volume, and with what quality? -
An expected mechanism of improvement
The assistant retrieves permitted sources, produces a cited synthesis, and reduces search and assembly effort. The analyst still reviews the result. -
A value translation
Saved analyst time becomes capacity available for higher-value analysis, stakeholder engagement, or additional decision support. It becomes cash savings only if spending or headcount can genuinely be reduced or avoided. -
A validation plan and guardrails
The pilot must show that time savings are real and that citation quality, authorization enforcement, and user trust remain acceptable.
The distinction between a baseline and a hypothesis is important:
- A baseline is observed before the AI-supported workflow exists.
- A hypothesis is an assumption about the incremental change the AI workflow may produce.
- Value realization is the operational or financial result that actually occurs after adoption.
Do not skip straight from “the assistant can answer questions” to “it will save millions.” The causal chain must be explicit: the system changes a concrete task; the task changes an operational outcome; the organization captures a business benefit.
For ROI purposes, this is an internal AI use case. It supports employee productivity and decision support rather than directly generating a customer sale. It should also be treated as a structured interaction for the MVP: although analysts may ask varied questions, the supported workflow, approved data sources, target users, and success measures are deliberately bounded.

Calculating the Return on Investment (ROI) of AI | AWS Cloud Financial Management
Read the relevant sections of AWS Cloud Financial Management’s article to see a practical sequence for classifying AI work, selecting a defensible metric, establishing a pre-AI baseline, and interpreting cost per outcome.
In “Classify your AI Use Cases,” read the classification guidance. Focus on why internal productivity use cases need a measurable outcome rather than a loose claim of efficiency. Then, in “Determine your Business Value Metric,” read the metric and baseline discussion. Notice the warning about selecting a measure that can be manipulated. Finally, in “Calculate the ROI,” read the worked example. The domain differs from this capstone, but its logic is transferable: isolate the incremental outcome beyond the baseline, then relate its value to cost.
Define the unit of work before measuring time
“Research time” is too broad to be a sound value metric. One analyst may include an informal question, while another may count only the final brief. If the unit of work changes between baseline and pilot, any claimed saving becomes unreliable.
For this MVP, define an eligible decision-ready evidence pack as:
A request from an authorized Strategic Operations user that requires locating material from approved enterprise sources, synthesizing relevant evidence for a stated decision topic, and producing a brief with source references for human review.
The baseline should include the full current effort needed to create that pack:
- clarifying the request when necessary;
- searching approved repositories and source systems;
- reading, comparing, extracting, and organizing evidence;
- preparing the written summary or briefing material;
- attaching or verifying references;
- responding to normal review questions and factual corrections.
It should exclude work that the assistant is not intended to affect, such as a director’s final judgment, a multi-day external consultation, or substantive strategic debate after the evidence pack is delivered.
This definition gives you a fair comparison. During the pilot, measure the same end-to-end unit of work. Do not measure only the few minutes spent interacting with the assistant while ignoring the time required to validate citations or correct an incomplete answer.
Select one primary value driver
The strongest initial financial driver is analyst capacity returned through reduced evidence-preparation time. It is close to the platform’s actual mechanism and can be measured without claiming that the tool independently caused a better investment decision.
A sensible metric hierarchy looks like this:
| Layer | Metric for the cited evidence assistant | Why it matters |
|---|---|---|
| Adoption | Share of eligible evidence-pack requests completed with the assistant | A technically capable system creates no value if intended users do not use it. |
| Operational outcome | Median end-to-end minutes per eligible evidence pack | Directly tests whether the workflow becomes faster. |
| Primary value measure | Verified analyst hours returned per month | Converts the operational improvement into usable organizational capacity. |
| Financial translation | Capacity value using a finance-validated fully loaded hourly rate | Enables an investment comparison without pretending all saved time is cash savings. |
| Quality guardrails | Citation completeness and correctness, rework rate, user-rated usefulness, unauthorized-access incidents | Prevents the team from optimizing speed by sacrificing trust, accuracy, or access control. |
The metrics need different roles. Time saved is not enough on its own, because users may not adopt the solution or may need extensive rework. Citations shown are not enough, because a citation can be irrelevant or inaccurate. And an apparent productivity gain is unacceptable if the platform retrieves a source the user was not authorized to access.
A practical non-negotiable guardrail is:
No unauthorized disclosure of enterprise information is acceptable, regardless of productivity benefit.
Establish a credible baseline
A baseline is not a single person’s recollection that “this normally takes about an hour.” It is a small but disciplined observation of current work.
For the pilot, collect baseline data for a representative set of eligible evidence packs across the major request types. For each item, record:
| Field | Example purpose |
|---|---|
| Request type and complexity | Separates a simple policy lookup from a multi-source portfolio briefing. |
| Start and finish time | Measures end-to-end effort using the agreed task boundary. |
| Active analyst effort | Distinguishes working time from waiting for another team or system. |
| Sources consulted | Helps explain unusually complex cases and establishes retrieval patterns. |
| Rework or review corrections | Captures quality cost, not just initial drafting speed. |
| Analyst role or fully loaded cost band | Supports later valuation without exposing unnecessary personal data. |
| Eligibility reason | Ensures the item genuinely fits the assistant’s intended scope. |
Use the median completion time as the headline baseline because a few unusually difficult requests can distort an average. Also keep the P90, the time below which 90% of cases are completed, because senior stakeholders will want to know whether the assistant helps only routine work or also improves the difficult tail of the workflow.
Segment the baseline when necessary. A short factual request and a multi-source decision brief should not be blended into one number if their work patterns are fundamentally different. You may begin with one well-defined category for the MVP, then broaden only after the measure is stable.
The central counterfactual is simple:
What would an analyst have spent to create this same qualifying evidence pack without the assistant?
That question protects against overstating AI attribution. If a request would have been solved through an existing search tool in five minutes, it should not be credited with a thirty-minute saving. Similarly, if analysts use the assistant only for the hardest cases, comparing those cases with the average baseline will understate the benefit. Compare like with like.
Calculate capacity value, then make assumptions visible
Define the following variables for a monthly estimate:
- : eligible evidence packs per month
- : share of eligible packs completed using the assistant
- : baseline median analyst minutes per pack
- : assisted median end-to-end minutes per qualifying pack
- : minutes saved per qualifying pack
- : realization rate for saved capacity
- : fully loaded analyst cost per hour
The time saving is:
The gross capacity returned each month is:
The capacity value adjusted for realistic realization is:
The realization rate deserves careful treatment. It acknowledges that not every saved minute becomes useful capacity. Analysts may experience fragmented time savings, use the time for unavoidable administration, or encounter lower volumes than expected. A conservative pilot case may use a rate such as 50% to 70%, with the rate documented as an assumption to validate.
Illustrative baseline hypothesis
The following numbers are examples, not facts about your organization. Replace them with measurements validated by the business product owner and Finance.
| Assumption | Illustrative value | Evidence needed |
|---|---|---|
| Eligible evidence packs per month, | 320 | Work-management records, request logs, or a short baseline study |
| Baseline median effort, | 75 minutes | Time study of current workflow |
| Target assisted effort, | 45 minutes | Pilot hypothesis, then observed assisted-task data |
| Assistant adoption, | 60% | Eligible-user pilot usage and workflow observation |
| Realization rate, | 70% | Product-owner agreement and evidence of reinvested capacity |
| Fully loaded analyst rate, | $90 per hour | Finance-approved labor-cost assumption |
First calculate the expected saving per qualifying pack:
Then calculate gross returned capacity:
At the stated labor rate, the gross monthly capacity value is:
Applying a 70% realization rate gives:
The annualized capacity value, assuming these conditions persist, is:
That leads to a usable initial hypothesis:
During the first sustained year of use, the authorization-aware cited evidence assistant is expected to support 60% of 320 eligible monthly evidence packs, reduce median end-to-end preparation effort from 75 to 45 minutes, and return approximately 96 gross analyst hours per month. At a 70% realization rate and a finance-validated fully loaded labor rate of $90 per hour, this represents an estimated annual capacity value of $72,576, provided citation-quality and authorization guardrails are met.
This is substantially more credible than saying “the AI will save analysts 40% of their time.” It states the population, unit of work, time horizon, mechanism, value conversion, and uncertainty.
Capacity value is not automatically cash savings
If analysts remain employed at the same cost, the organization has not necessarily reduced expenditure by $72,576. It has gained capacity. That is valuable if leaders commit to reinvesting the time in work that matters, for example:
- increased coverage of strategic requests without adding headcount;
- deeper analysis of exceptions and emerging risks;
- faster preparation for decision forums;
- more time for stakeholder alignment and quality review.
Call this capacity value unless the business can point to a cashable mechanism: avoided contractor spend, avoided hiring, a reduced service cost, or a committed redeployment that creates separately measured revenue or margin. This distinction is especially important in executive investment cases, where overstated “labor savings” quickly erode trust.
Also avoid double-counting. If you claim the value of analyst hours returned, do not add a second dollar value for “faster decisions” unless you can isolate an additional, evidence-backed financial effect.
Include costs before using the word ROI
The annual capacity-value estimate is the benefit side of the model. A decision case also requires the full incremental cost of delivering and operating the capability.
For this platform, costs typically include more than model calls and cloud infrastructure:
| Cost category | Examples for this MVP |
|---|---|
| One-time implementation | Discovery, engineering, source onboarding, data preparation, evaluation-set creation, security design, and pilot setup |
| Change and enablement | User onboarding, training, workflow documentation, communications, and support preparation |
| Recurring technology | Model inference, embeddings, retrieval and storage, hosting, logging, monitoring, and security services |
| Recurring operating effort | Source curation, access review, evaluation, incident handling, prompt or configuration maintenance, and product support |
| Governance and assurance | Privacy, legal, security, records, model-risk review, and audit evidence |
How to Create a Business Case for Your Digital Transformation [ROI and Investment Case Analysis 101]
Watch selected segments of “How to Create a Business Case for Your Digital Transformation” by Digital Transformation with Eric Kimberling. The video is useful here because it distinguishes visible software costs from the broader implementation and change costs that make up total ownership, then frames the business case as an instrument for realizing benefits.
Watch full ownership for the discussion of direct and indirect costs, including internal labor, data work, infrastructure, and change management. Then watch measurable benefits to contrast concrete, attributable benefits with broader business outcomes. Finish with benefits realization, focusing on the need to assign ownership and operational changes behind a high-level target.
Suppose, for illustration, that the MVP has a $24,000 one-time delivery and onboarding cost, plus $2,200 in monthly operating cost during its first year:
Using the earlier capacity-value estimate:
If the organization and Finance accept capacity value as the appropriate benefit basis, the corresponding first-year ROI estimate is:
This is a scenario, not a forecast. It assumes stable volume, sustained adoption, the target time saving, a valid 70% realization rate, and no guardrail failure requiring a redesign. State those assumptions on the same page as the percentage.
A useful unit-level companion metric is cost per qualifying evidence pack. If 60% of 320 monthly eligible packs use the assistant, that is 192 assisted packs per month. With $2,200 of monthly run cost:
This metric helps monitor scaling. It can improve as usage rises and fixed operating costs are spread across more qualifying work, but it can also worsen if complex prompts, expensive models, or low adoption increase cost faster than value.
Make the hypothesis govern the pilot
A business case should guide choices during delivery, not become a forgotten approval slide. Turn the hypothesis into a measurement contract between the product owner, delivery team, Finance, and governance stakeholders.
Your first version should contain the following:
| Element | What to record |
|---|---|
| Use-case boundary | Authorized users, approved sources, request types, and activities outside scope |
| Baseline period | Dates, number of representative work items, and method used to collect time and quality data |
| Benefit assumptions | Volume, adoption, baseline time, target assisted time, hourly rate, and realization rate |
| Cost assumptions | One-time costs, monthly run costs, owner for each estimate, and cost-allocation method |
| Primary success measure | Verified monthly analyst hours returned and capacity value |
| Guardrails | Authorization incidents, citation-quality criteria, rework, user usefulness, and escalation conditions |
| Review cadence | Weekly operational review during pilot; formal value and risk review at the pilot decision gate |
| Decision owner | Executive sponsor for scale, revise, or stop; product owner responsible for benefits realization |
Measure post-pilot results against the baseline using the same task definition. Keep a record of changing conditions, such as a newly onboarded source, different model provider, expanded user group, or revised approval process. These changes may improve the system, but they also make before-and-after comparisons less clean.
The result should be presented in three scenarios rather than one precise-looking number:
- Conservative: lower adoption, smaller time reduction, and lower realization.
- Expected: the assumptions currently supported by discovery and stakeholder input.
- Upside: higher sustained adoption and a demonstrated ability to reinvest capacity.
The executive decision is not “Which number is correct?” It is whether the expected value, uncertainty, costs, risks, and strategic learning justify a bounded pilot.
Key takeaways
A baseline value hypothesis converts an attractive AI idea into a claim that can be tested and governed.
For the cited evidence assistant:
- classify it as an internal, bounded decision-support use case;
- define a consistent unit of work: the eligible decision-ready evidence pack;
- establish a pre-AI baseline for volume, end-to-end analyst effort, rework, and quality;
- estimate value from eligible volume, adoption, verified minutes saved, a realistic realization rate, and a Finance-validated labor-cost assumption;
- describe saved time as capacity value unless there is a clear route to cash savings;
- include delivery, data, governance, change, and ongoing operating costs before calculating ROI;
- keep citation quality and authorization enforcement as guardrails, not tradeable benefits.
Your capstone artifact from this lesson is a one-page value hypothesis with its assumptions, formulas, cost categories, guardrails, owners, and pilot measurement plan.
Next, you will create a prioritized register of technical, organizational, and responsible-AI risks. That register will test the conditions under which this value hypothesis is allowed to hold—and identify the risks that could reduce, delay, or invalidate it.
Can't find a good explanation? Sign up and we'll make it for you
Sign up