ANAlpesh Nakrani
SolutionsBlogBooksPraiseAboutWork with me
Back to the blog
Blog/Sep 14, 2026 · 10 min

How to Calculate AI ROI Before Choosing a Model or Vendor

AI ROI is workflow economics: cost per unit of work before and after AI, at real volume, minus what errors cost to catch and fix.

AI ROI is not a model score or a price-per-token comparison. It is workflow economics: what a unit of work costs today, what it costs with AI in the loop at real volume, and what you lose when the AI is wrong and a human has to redo it. Skip that calculation and you are choosing a vendor on vibes, no matter how many benchmarks are in the deck.

Teresa Kwan is a composite I run into constantly: CFO at a 400-person accounts payable services firm processing roughly 60,000 vendor invoices a month for mid-market clients. Three AI vendors had pitched her team on invoice coding and matching automation. Each deck opened with a benchmark score and a per-invoice API cost. None of them opened with what she actually needed: the current fully-loaded cost of coding an invoice by hand, the volume that cost scales against, and what it costs when the model miscodes an invoice and nobody catches it until reconciliation three weeks later.

Teresa's team spent two weeks building that baseline before they evaluated a single vendor. It slowed the decision down. It also showed that the cheapest vendor by per-invoice API cost was the most expensive vendor once its higher miscoding rate was priced against their actual rework cost, a number nobody had calculated before because nobody had needed to.

I've spent fourteen years in this industry, IC to CTO to COO to VP of Growth at ViitorCloud, and the buyers I talk to who get burned by an AI purchase almost never got burned by a bad model. They got burned by a vendor comparison that never had their own numbers in it.

Key takeaways

  • AI ROI is a workflow economics problem, not a model comparison. The number that matters is cost per unit of work before AI versus after, at real production volume.
  • Benchmark scores and per-token pricing tell you almost nothing about payback. They ignore volume, integration cost, and what a wrong answer costs to fix.
  • Error and rework cost is usually the biggest hidden variable, and most companies have never measured it because nobody needed to before AI made the question urgent.
  • A clean ROI model requires baseline data most companies haven't instrumented, which means the "ROI calculation" often has to start as an instrumentation project, not a spreadsheet.
  • Vendor selection should follow the model, not precede it. Decide what workflow economics you need first; let that decide which model and vendor clear the bar.

Why the usual approach to AI ROI fails

Most AI buying decisions start backwards. A vendor or a model gets shortlisted first, usually on a benchmark score, a demo, or a price-per-million-tokens comparison, and the ROI case gets written afterward to justify a decision that has effectively already been made. That ordering guarantees a weak business case, because the inputs that actually determine payback, baseline cost, volume, and error rate, never made it into the comparison in the first place.

This is close to the same failure I've written about in why most AI pilots never become production revenue: the technical work gets done well and the commercial case gets skipped entirely. With ROI specifically, the skipped step is even more basic. Most finance and operations leaders can tell you what a vendor charges per API call. Few can tell you, to the dollar, what a unit of work costs today, fully loaded, including the rework nobody tracks as its own line item.

A benchmark score tells you how a model performs on someone else's task. It tells you nothing about what a unit of your work costs before and after you deploy it.

The workflow economics framework

The model I use with clients has four inputs, and all four have to be real numbers pulled from your own operation, not vendor marketing or an industry average.

  • Baseline cost per unit of work. Fully loaded labor cost, including overhead, to complete one unit today: one invoice coded, one ticket resolved, one claim adjudicated. Most companies have a rough version of this and almost never have it broken out by task step.
  • Volume at production scale. Not the volume in the pilot. The volume the workflow actually runs at in a normal month, because unit economics that look fine at 500 units a month often collapse, or improve dramatically, at 60,000.
  • AI cost per unit at that volume. Inference cost, licensing, and integration and maintenance overhead, amortized per unit at real volume, not the sticker price in the vendor's pitch deck.
  • Error and rework cost. The cost of catching and fixing a wrong output, multiplied by the rate at which the AI produces one, compared honestly against the error rate of the process you're replacing.

Put those four together and the arithmetic is simple, even when the inputs are hard to get. Net savings per unit equals baseline cost minus AI cost minus the product of AI error rate and rework cost, compared against the same calculation for the human-run process it replaces. Multiply by volume and you have an actual payback period, not a vendor's promised one.

In Teresa's case, the fully-loaded cost of coding an invoice by hand ran close to $4.20, with a human error rate requiring downstream correction of about 2%. The cheapest AI vendor priced out at $0.38 per invoice, but its miscoding rate on unfamiliar vendor formats ran closer to 11%, and every miscoded invoice cost roughly $18 to catch and fix during reconciliation. At 60,000 invoices a month, that vendor's true cost per invoice, once rework was priced in, came out higher than a competitor charging nearly triple the per-invoice rate with a miscoding rate under 3%.

The honest trade-off: most companies haven't instrumented for this

Here is the part vendors don't put in the deck and consultants don't like to say out loud: most companies cannot run this model on day one, because they have never measured the inputs it needs. Fully-loaded cost per unit of work, broken out by task, is not a number most finance systems produce automatically. Error and rework cost is even rarer, because until AI raised the question, nobody needed to price the cost of catching a mistake as its own line item.

That means the real first step in an AI ROI calculation is often an instrumentation project: a few weeks of time-and-motion measurement, a review of how the team currently catches and fixes errors, and a clear-eyed baseline before a single vendor gets evaluated. That is slower than picking the vendor with the best demo. It is also the only way the ROI number you eventually present to your board survives contact with a real budget cycle, instead of getting quietly revised downward in year two.

The instrumentation project is not overhead standing between you and the AI decision. It is the AI decision. Everything after it is arithmetic.

What the data says about the gap between adoption and payback

The instrumentation gap shows up clearly in the research. McKinsey's most recent global AI survey found that 88% of organizations now use AI regularly, but only 39% report any measurable enterprise-level financial impact, and just 6% qualify as high performers seeing more than 5% EBIT impact from it. Adoption and payback are not the same curve, and most companies are only tracking the first one.

A Microsoft-sponsored IDC study puts a number on the dispersion behind that gap. It found that organizations deploying generative AI see an average return of $3.70 for every dollar invested, but leaders in the study saw $10.30, roughly triple. An average that wide is not describing one distribution. It is describing two different populations: companies that modeled the workflow before they bought, and companies that didn't.

BCG's research on the same question points to the mechanism behind that split. Its analysis found that only about 5% of companies are generating significant value from AI at scale, while nearly 60% report little or no impact, and the difference was rarely the model. It was whether the company redesigned the workflow the model sits inside, rather than bolting AI onto the existing process and hoping the economics worked out.

The ViitorCloud view: model the workflow before you shortlist a vendor

Most AI vendor conversations I sit in on start with a capability question: what can this model do. The question that actually predicts whether the deal pays back is a finance question: what does a unit of this workflow cost today, at what volume, and what happens to that number when the model is right 95% of the time instead of 100%. Get that number wrong, or skip it, and the vendor selection that follows is decoration on a decision you can't defend at renewal.

That is the gap our AI ROI Calculator and Workshop is built to close: a short, structured engagement that builds the four-input workflow economics model with your own numbers before you sign anything, so the vendor comparison that follows is measuring the right thing. It runs through ViitorCloud's technology consulting practice, and it plugs into the same unit-economics discipline I've written about on the inference side: the cheapest model that clears your accuracy bar usually wins the workflow, once error and rework cost are actually in the model instead of assumed away.

The honest trade-off is time. A real workflow economics model takes longer to build than reading a vendor's ROI calculator off their own website, and it will occasionally tell you the AI project doesn't pay back yet at your current volume or error tolerance. That is not a failure of the exercise. Killing a bad bet before the contract is signed is cheaper than unwinding it after.

The AI ROI checklist

Run this before you shortlist a vendor, not after.

  • You have a fully-loaded, per-unit cost for the workflow today, broken out by task step, not a department-level average.
  • You are modeling AI cost at real production volume, not the volume in the pilot or the vendor's demo.
  • You have measured, or are actively measuring, your current process's error and rework rate before comparing it to the AI's.
  • You have priced what an AI error costs to catch and fix, not just how often it happens.
  • The payback period you present is calculated from your own numbers, not adapted from a vendor's case study or an industry average.

Frequently asked questions

How do you actually calculate ROI on an AI investment?

Model it as workflow economics: baseline cost per unit of work today, AI cost per unit at real production volume, and the cost of catching and fixing AI errors multiplied by how often they happen. Compare that net cost to the process you're replacing, at your actual volume, not a vendor's demo volume.

Why do benchmark scores and per-token pricing fail to predict AI ROI?

Because they leave out the three variables that determine payback: your actual volume, your integration and maintenance cost, and what a wrong answer costs to fix downstream. A model can win every public benchmark and still lose the ROI calculation once error and rework cost are priced in at your real volume.

What if we don't have the baseline data to build this model?

Then the first project is instrumentation, not vendor selection. A few weeks of measuring current cost per unit of work and current error and rework rate is slower than picking a vendor off a demo, but it's the only way the ROI number survives a second budget cycle instead of getting quietly revised downward.

Should the CFO or the CTO own the AI ROI decision?

Both, because the inputs come from both seats. The CTO or CIO owns the AI cost and error-rate side of the model; the CFO or COO owns the baseline cost and rework-cost side. An ROI model built by only one side is missing half its inputs.

Share
Next

Keep reading

View all blogs

Ask AI about How to Calculate AI ROI Before Choosing a Model or Vendor