Your model bill keeps climbing, you still can't tell if the output is worth it, and the number everyone is arguing about was never the one that decides AI ROI.
Satya Nadella called himself a tokenmaxxer. "It's addictive," he said, and costly.
He is right on both counts. He is also reading the wrong number off the page.
Tokenmaxxing showed up in the last year as a name for burning as many AI tokens as you can in pursuit of more output. The bills got attention fast. Disney started warning its engineers against it. Synthesia's HR chief argued that managers should stop counting tokens at all. Grow them, cut them, or stop counting them: every camp is still arguing about the tokens. The loudest version asks how cheap we can make AI. It feels like the responsible question. It is also the wrong place to start a fight about AI ROI.
You feel the pull every time the invoice arrives. You wired AI into your work, the bill came in heavier than last month, and when you look at what it produced you can't honestly say whether you bought leverage or just bought motion.
Cost is real. It is also one term in a product of several, and almost never the term that decides how the month went.
A token count is seductive because it is precise. It updates in real time, it fits on a dashboard, it ranks people on a leaderboard. Precision feels like insight.
It rarely is. A number that measures how much you spent tells you nothing about what you got back. You can burn a fortune in tokens producing code nobody ships, drafts nobody reads, analyses nobody acts on. The dashboard will look industrious the entire time.
Any single zero cancels the whole equation
Strip AI ROI down to what actually moves it and you get four things that multiply.
Value created (V) is human capability (H) times AI capability and task fit (A) times the economic leverage of the problem (L) times your ability to capture the result (C): V = H × A × L × C. Call the cost K, your hours plus the model's bill. Then your return is what the work was worth minus what it cost, over what it cost: AI ROI = (V − K) / K.
The shape matters more than the symbols. Because this is multiplication, every term can veto the others. Push any one of the four to zero and the whole product collapses, no matter how large the rest are. A brilliant engineer (high H) with a capable model (high A) who captures the value cleanly (high C) still produces nothing if the problem doesn't matter (L near zero).
The mistake underneath the cost panic is quieter than it looks: we count work as if it adds up. More tokens, more output, more tickets, one running total you grow by feeding it. Almost nothing that creates value behaves that way. It multiplies, and a multiplied system has no safe place to be weak.
Cheap tokens cannot rescue a zero. They can only make the zero arrive faster and at a lower unit price.
You don't plug real numbers into this. You use it to find which of the four you are actually short on.
Why the human is usually the biggest number in AI ROI
Of the four terms, human capability tends to swing the hardest, which is the part the cost conversation skips entirely.
A capable person changes the equation before the model ever runs. They pick a problem worth solving. They hand the model the context that makes its answer good. They recognize a wrong answer in seconds instead of shipping it. They take a rough output and turn it into something that survives in the real world.
Hand the same model and the same token budget to someone without that judgment and you get the inverse. METR ran a controlled study and found experienced developers took longer with AI tools, even while they believed the tools had sped them up.¹ The model didn't get worse between the two people. The human did.
This is why an excellent engineer can earn a real return on an average model with expensive tokens, as long as the problem is worth it. The expensive ingredient here is attention. Tokens were always the cheap part.
Running the equation on yourself has a catch. Three of the four terms you can check as you go: the model's fit you read off its output, the leverage you argue on the merits, capture you learn the moment you try to ship. The fourth, your own judgment, you mostly grade by feel, and feel is the one signal a fluent answer corrupts. A confident wrong answer reads just like a confident right one. Results will correct you, and so will a sharp reviewer, but you've already acted on the feeling by then. The METR developers were sure they'd sped up while the stopwatch said the opposite. Trust the outcome, not the certainty.
A great engineer on a trivial problem still loses
Human capability alone is the opposite trap. You need it, and on its own it gets you nothing.
Point your best person and your best model at a problem nobody needed solved and you have spent premium attention to produce something elegant and worthless. Leverage is the term that decides whether the value is even large enough to bother with.
Capture is the term that decides whether you keep any of it. You can produce genuinely valuable output and still see no return, because you can't ship it, can't sell it, can't get a single user to adopt it, or you watch a competitor take the upside while you eat the cost. The value existed. It just never landed in your account.
So "hire great people and let them tokenmaxx" is half a sentence. The whole sentence is great people, on high-leverage problems, whose value you can actually capture.
And if you're building solo, there's no "them" in that sentence. You're the judgment, the problem pick, and the capture, all at once, and a weak term in any of them is yours alone to carry. Nobody on the team covers it, because there's no team. The equation gets less forgiving the smaller you are.
Output doesn't count until it survives
Here is the move the whole industry keeps fumbling: treating output as if it were value.
More code, more drafts, more tickets closed, more tokens burned. All of it looks like productivity. It is the same error as ranking photographers by shutter count. The one who fired four thousand frames did not out-shoot the one who took forty and printed two that mattered.
This isn't a quirk of token dashboards. It's what happens to any craft the moment someone finds a number to rank it by.
None of it is worth anything until it survives contact with reality. Validated output is output with evidence behind it, and the evidence climbs in levels:
- Technical: It works. The code runs, the draft is coherent, the analysis holds.
- Behavioral: Someone uses it. A real user, not the person who generated it.
- Outcome: It moves a metric. Conversion, retention, cycle time, or cost, something measurable shifts.
- Economic: It shows up in the money. The better outcome creates profit or reduces cost, in dollars you can point to. (Revenue is a tempting stand-in here, and a weak one, because revenue can climb while profit stays flat.)
Raw output sits below the first level. A token count sits below even that. It measures what you spent. The distance you actually moved is a different number entirely. When you make the spending gauge the goal, you get Goodhart's law on schedule: people optimize the number and quietly abandon the thing the number was supposed to stand for.²
Put that number on a leaderboard and the failure compounds: you reward the people who move it. Motion ranks, and the quieter operator who solves the thing cheaply and ships something that survives barely registers on the chart.
Economic validation, the top level, is the real endpoint. It is also delayed and hard to attribute, which is why nobody can ever quite prove the profit line moved because of the model. So watch a more immediate measure instead:
AI leverage = expected economic value of validated output ÷ (human time cost + AI cost). The honest unit underneath it is validated economic value per scarce human hour: the worth of output that survived the levels above, over what it took to produce. Watch that ratio move across a quarter and you are measuring something real, long before the profit line catches up.
When burning more tokens is the rational move
Hold value capture steady, assume you can ship what you make, and run the rest of the combinations. The verdict stops tracking the cost column.
Human judgment | Model fit | Problem leverage | Token spend | Likely AI ROI |
|---|---|---|---|---|
Strong | Good | High | Costly | Strong improvement |
Strong | Poor | High | Costly | Modest to strong |
Strong | Good | High | Cheap | Strong improvement |
Strong | Good | Low | Cheap | Little effect |
Weak | Good | High | Cheap | Long-term worsening |
Weak | Good | Low | Cheap | Long-term worsening |
Weak | Poor | High | Costly | Worsening |
Read straight down the token-spend column. It flips between cheap and costly while the verdict barely moves. Read the human-judgment column and the verdict turns with it every time: strong judgment keeps you out of the red, weak judgment drops you into it no matter what else is true. Leverage and model fit decide how far above the line you land. Cost decides almost nothing.
Which hands you the actual rule. When the human is strong, the model is useful, the problem has real leverage, and you can capture what you create, spending more tokens is rational. Burn them. The return is sitting right there, and the tokens are the cheapest thing standing between you and it. Optimizing that cost first is a false economy. You save pennies on the one input that was never the constraint.
When the human is weak, or the problem is trivial, or you can't capture the value, no token budget saves you. Cheap AI is worst of all in that case, because it lets you manufacture plausible, confident, wrong work faster than anyone can review it. You don't save money. You finance a larger cleanup, and the technical debt it leaves behind.
The scarce resource was never the tokens
Cheap AI does not save weak judgment. Expensive AI does not ruin strong judgment aimed at a problem that matters.
The bottleneck moved while everyone was staring at the invoice. It sits in the human now: in choosing the problem, in feeding the model, in catching its confident mistakes, in turning a draft into something that ships and sells. Tokens are abundant and getting cheaper every quarter.
Capable attention pointed at leverage is the thing in short supply. You cannot tokenmaxx your way out of being short on that.
Rabbit Hole:
When AI gains keep evaporating before they reach the bottom line, Promised 10x, Got 2x. Why, and How to Fix It is about why the multiplier shrinks on the way down.
For why the human term dominates the equation, What AI Actually Searches When It Helps You Think makes the case that your own knowledge sets the ceiling on what a model can do for you.
And on the capture term, Pay Per Result Might Be the Unit Test for Pricing AI SaaS looks at tying price to outcomes instead of usage.
Footnotes:
- METR, a research group, ran a randomized trial on AI coding tools in 2025 and found experienced developers took longer to finish tasks with AI assistance, even though they predicted and felt the opposite. The gap between felt speed and real speed is exactly where AI ROI quietly leaks.
- Goodhart's law: when a measure becomes a target, it stops being a good measure. The Wikipedia entry on tokenmaxxing cites it directly, because a token count is a textbook case of optimizing the proxy and losing the thing the proxy stood for.