Practical PlaybookFINOPS
The Cheaper AI Model Can Produce the More Expensive Answer
- Author
- Alex Florian
- Published
- Updated
- Reading time
- 4 min
The model bill goes down, but reviewers spend longer fixing the answers. Both observations can be true: the team bought a cheaper input and may now be running a more expensive service.
Take an assistant that drafts support replies. After a switch to a lower-priced model, more drafts need another attempt or substantial editing before an employee can use them. The dashboard records the model saving; the review team absorbs the extra work.
The question behind unit economics is simple enough to ask without a finance formula: what does it cost to obtain a result we can actually use? The provider's billing unit may be tokens—the pieces of text processed or generated—but the support manager needs an acceptable reply, not a collection of inexpensive tokens. [1][2]
Give “accepted” a concrete meaning
For this drafting workflow, an accepted result could be a reply that a reviewer can approve under the service's agreed accuracy, relevance, and safety requirements. The review work required to get there belongs in the comparison.
That unit is narrower than a resolved support case. A good draft may still require delivery, a customer response, or another team action before the case is complete. Cost per accepted draft and cost per resolved case answer different questions; neither should inherit the other's claim by sharing the word “outcome.”
The definition needs to remain stable. If a new model appears cheaper because reviewers start accepting unfinished answers, the ratio has changed partly because the standard changed. Counting every generated paragraph as useful work would conceal the reason the assistant was introduced.
The more expensive design can produce cheaper usable drafts
Consider two designs tested on the same representative set of 1,000 drafting requests over a comparable period. These are hypothetical teaching figures, not provider prices or client results. Both use the same acceptance rule and the same defined drafting-cost boundary, including unsuccessful attempts and the review counted within that workflow.
| Illustrative design | Cost across all relevant attempts | Accepted drafts | Cost per accepted draft |
|---|---|---|---|
| A | $1,000 | 500 | $2.00 |
| B | $1,200 | 800 | $1.50 |
Design B spends $200 more overall but produces accepted drafts at a lower unit cost: $1,200 divided by 800 is $1.50. Design A's $1,000 divided by 500 is $2.00.
The calculation doesn't establish the full cost of resolving every support case. Requests without an accepted draft still need an appropriate route, and that wider service cost must be included before making a wider claim. It does show why a smaller model bill is insufficient evidence for choosing the cheaper drafting design.
Failed attempts matter because they are part of obtaining the successes. Removing their costs would make the accepted drafts look cheaper without changing the actual work. The same happens when correction effort is excluded because it sits in another team's budget. [1]
Compare the costs that could change the choice
The relevant boundary may include model use, hosting, retrieval of source material, retries, and human review. Not every small experiment requires perfect allocation of every shared expense, but a plausible omitted cost deserves attention when it could reverse the recommendation.
For the original model switch, I would begin with the work the reviewers say increased. Are they correcting more drafts? Spending longer on a particular type of request? Sending more work back for another attempt? That comparison can explain whether the lower input price survives the journey to an accepted answer.
The average can hide important differences. Routine replies and difficult policy questions may have different economics, especially across languages or customer groups. A model can be a sensible choice for one supported class without being the best route for every request.
A material privacy or safety failure also needs separate consideration. Enough inexpensive successes cannot make an unacceptable action disappear inside an average. The quality and safety conditions are part of what makes the result usable, not optional additions after the ratio is calculated.
Let the comparison reach the people who can change the workflow
A unit-cost result is valuable when the model team and service owner can use it together. They might keep the lower-priced model for suitable requests, choose another model for a difficult class, or fix the surrounding retrieval and review process. The evidence should identify which intervention explains the difference, rather than crediting the component with the most visible price.
For this decision, the reviewers' extra effort is not a footnote to an engineering saving. It is part of producing the answer. Once it is counted within the appropriate boundary, the team can choose a design for the work it delivers—not merely for the price printed beside the model name.