All insights

AI in the Organization

The hours saved are not yet the return

Trace AI value from real use to changed workflow behavior, operational results, and financial outcomes with a practical measurement card.

An evidence chain connecting AI use, changed work, operational results, and financial results

An anonymized US eyewear wholesaler told us that its tailored ERP and AI-assisted reporting workflow gave the team about ten hours back each week, which is a useful result and also the opening of the ROI question.

Perhaps reports arrived earlier, the team investigated more exceptions, overtime fell, or the business handled more orders without hiring, but we don’t have evidence for any of those outcomes, so we shouldn’t quietly convert the estimate into payroll savings. Maybe the honest claim is smaller (and still worthwhile), because a recurring task consumed less of the team’s week. Even so, the figure was the client’s estimate, not an instrumented time study, and the ERP mattered alongside the AI-assisted part.

That may sound overly cautious, but the next measurement depends on what the business hoped to do with the returned capacity, and if the original problem was an overdue report, we should look at when the report now arrives and whether anybody can act sooner. If the team was overwhelmed, we’d look at backlog, overtime, or completed orders, because multiplying ten hours by an hourly salary would skip the part where time becomes valuable (and sometimes it never does).

Another anonymized project makes the limitation clearer. A US presentation-training business used an AI coaching product to open a new revenue line, so hours saved would be the wrong outcome there and the useful questions concern customer uptake, coaching quality, retention, delivery cost, and the gross margin of an offer that didn’t exist before. Two projects can both contain AI while having entirely different economics, which is one reason we prefer to look at the queue or outcome that needs to change instead of beginning from a general promise of efficiency.

AI creates a return only when use changes the work and that change reaches a financial resultAn evidence chain connects actual use to changed behavior, an operational result, and finally financial value.THE EVIDENCE CHAIN01USE02BEHAVIOR03OPERATION04FINANCEA return exists only when every link can be evidenced
AI creates a return only when use changes the work and that change reaches a financial result

Follow the result as far as the evidence goes

Before reaching for an ROI spreadsheet, we’d follow a simple chain through the work. First we look for repeated use on eligible cases, and then for changed behaviour (fewer handoffs, a different review pattern, perhaps new ownership of exceptions). Next comes the operating result, which could be cycle time, throughput, rework, backlog, or quality, and only after that do we get to overtime, avoided cost, gross margin, conversion, retention, or another financial consequence.

Nevertheless, the chain helps because usage can be surprisingly uninformative. A person may send dozens of prompts while doing the task twice, once with AI and once manually to check it, because drafting may become faster while review grows enough to make the complete workflow slower. Even an operational improvement isn’t automatically cash (extra capacity has financial value only when the business can use it).

OpenAI’s deployment guidance makes a similar distinction between activity and impact. Its measurement guide connects bottom-up efficiency to results such as cycle time, customer satisfaction, and revenue capacity, while its evidence guidance treats repeat use and a changed team process as stronger evidence of adoption than a demonstration or one trial. These are vendor guides, not independent proof, but the distinction is practical.

Therefore, we’d choose one primary kind of value before the pilot, since capacity means more acceptable work or a smaller backlog only when demand exists, while quality needs a defined rubric and must include errors the system introduces. Revenue should be judged through contribution margin (not a headline sales number), and avoided cost needs an observable expense such as overtime, contractor work, or rework. An “avoided hire” is credible only when workload and hiring intent support it. A risk project may begin with a control result instead, perhaps that every exception is now recorded and reviewed on time.

Save a small baseline before it disappears

Once the new workflow becomes normal, reconstructing the old one is surprisingly hard, so before the pilot, we’d save a card naming the workflow, one unit of work (a report, case, order, or request), the eligible population, the primary outcome, and any guardrails. It would include the sample period and number of cases, elapsed and human touch time, acceptance without rework, important exceptions, backlog, and cost per acceptable case. Finally, we’d note the evidence source, owner, and any known seasonal or selection effect. It isn’t a grand measurement system, but it’s the small piece of paper that stops everybody remembering the old process differently three months later.

Moreover, the unit of work matters more than it first appears, because “hours saved per employee” mixes unrelated tasks, whereas “minutes of review per accepted claim file” can be observed, and we also need the denominator because ten corrected outputs out of 100 attempts tells a different story from ten out of 10,000, and an average can hide the one failure that consumed a day of repair.

Consequently, we measure the complete task. Tool cost, setup, prompting, review, correction, exceptions, retries, integration, monitoring, training, and downstream rework all belong somewhere in the account, because a cheap model that creates more retries may cost more in practice and we care about cost per acceptable completed unit, not cost per call.

For a concrete way to test acceptable completion, review effort, dangerous failures, and release boundaries together, use the end-to-end AI workflow evaluation method.

NIST’s AI Risk Management Framework recommends documenting test sets, metrics and evaluation methods. That discipline isn’t only for model safety, but makes the business case auditable (someone else can see what was tested, on which cases, and under which definition of success).

A published result isn’t our forecast

In a field study of 5,179 customer-support agents, Brynjolfsson, Li and Raymond found that a generative AI assistant increased issues resolved per hour by nearly 14% on average in one customer-support deployment. The gains were much larger for less experienced workers and small or negative for the most experienced group.

Furthermore, the variation is the useful part, because this is good evidence from one production setting and not a percentage to paste into a proposal for wholesale reporting, clinical intake, or sales operations. Our baseline, users, task mix, quality threshold, and adoption determine our result. A pilot made only of easy cases or enthusiastic volunteers may hide the cost of exceptions and broad rollout, while a simple before-and-after comparison can confuse the new system with seasonality, staffing, or a policy change.

An SME doesn’t always need an elaborate controlled study (especially for a low-risk decision). However, we do want representative cases, the same definition of acceptable work on both sides, a note of other changes during the test, and a clear label separating measured data from employee estimates. When attribution or future demand is uncertain, a range is more believable than a precise forecast.

Only then would we translate the portion with a real economic consequence, and the calculation can stay plain because net value is attributable financial value minus build, integration, migration, training, operation, review, maintenance, and change costs. It’s easy arithmetic once we’ve earned the values on either side (earning them is the difficult part, of course).

We shouldn’t turn all returned time into salary savings unless payroll actually changes, because the stronger case may be that a constrained specialist serves more customers, a report arrives before a decision, or the team absorbs growth without degrading service.

The OECD notes that intangible investments such as training and changes to business processes and software complicate productivity measurement. Those costs are particularly easy to forget because they aren’t printed on a model invoice. We also wouldn’t value every returned hour at salary unless payroll actually changed. The better case may be that a constrained specialist can serve more customers, the report arrives before a decision, or the team absorbs growth without poorer service.

Stop the sentence where the evidence stops

Before the test, we’d agree that the result can lead us to continue, change, narrow, or stop. A negative result may reveal a data problem, unclear ownership, an obsolete approval, or simply no demand, which is still useful evidence provided we don’t rescue the launch by choosing a new outcome after seeing the old one fail.

The most credible ROI sentence is sometimes modest:

The team used the workflow on 140 eligible cases. Median preparation time fell, reviewer effort stayed level, and the error rate remained within the agreed threshold. We haven’t yet shown an effect on conversion.

That sentence gives a leader something to decide with, and it tells us what to measure next. Time saved can be the outcome when overtime falls or a backlog clears, it can be a path to value when the capacity is used for more productive work, and it can disappear somewhere between a survey and the income statement. We don’t need to make the result sound larger. We need to stop the claim at the last link the evidence supports.

Common questions

Before you decide.

How do you calculate ROI for an AI project?

Compare the financial value attributable to the changed workflow with its full cost over the same period. Include implementation, integration, review, training, operation, errors, and maintenance rather than only model or licence fees.

Is time saved a valid AI ROI metric?

It is valid evidence of efficiency, but it becomes financial value only when the capacity is used: less overtime, avoided hiring, more completed work, faster revenue, better service, or more expert attention.

What should an AI pilot measure?

Measure adoption in real work, successful task completion, quality and exception rates, elapsed workflow time, human review effort, operating cost, and the specific business outcome the workflow was meant to change.

How long should an AI ROI study run?

Long enough to include a representative number of real cases and normal variation in the workflow. The right duration depends on task frequency, seasonality, learning effects, and the consequence of errors.