According to a report published by The Wall Street Journal on October 5, 2026, only 11% of nearly 400 surveyed companies can accurately predict their AI spending in advance. The remaining 89% — nearly nine in ten companies — cannot accurately forecast how much they will spend on AI services. The publication points to the reason: AI models' token consumption and the final bill vary sharply from prompt to prompt. The survey was conducted recently, and its results were published in WSJ's technology section.

To cut costs, companies typically choose models with lower list prices, reasoning that a lower token price means a lower total bill. But new research by scientists at Stanford University, Carnegie Mellon University, UC Berkeley, and Microsoft Research shows this logic does not always hold: in some cases the cheaper model ends up costing more.

How the Study Was Conducted

Scientists ran models from different price tiers through more than 6,800 tasks. In each task, models were compared pairwise: a cheaper and a pricier model performed the same prompt, after which their total token consumption and final cost were compared. Across these pairwise comparisons, in 32% of cases the total cost of the model with the lower list price exceeded that of the higher-priced model.

The researchers called the phenomenon the "price reversal". In some pairs the reversal gap reached up to 28x — in the most extreme cases, the final bill of the model chosen as the cheaper one came out 28 times higher than that of the pricier model. The researchers emphasize that the list price shows only the price of a single token; the final bill depends on how many tokens the model spends in total on a given task. So a cheap token does not always mean a cheap final bill.

The paper, titled "The Price Reversal Phenomenon: When Cheaper Reasoning Models Cost More", was published on the arXiv preprint server on May 28, 2026, under number 2603.23971v2. WSJ's October 5, 2026 report drew on the findings of this very study. The authors are Lingjiao Chen and colleagues from Stanford University, UC Berkeley, Carnegie Mellon University, and Microsoft Research.

One Prompt, Two Bills

The practical example cited by WSJ makes the gap clear. The pricier Gemini 3.1 Pro model completed the given prompt in 85 steps, spending $1 in total. When the same prompt was given to the cheaper Gemini 3 Flash model, it took nearly 1,000 steps — almost 12 times more in step count — and ultimately failed to complete the task. Even so, the tokens consumed cost $14, and that amount stayed on the bill.

In other words, the cheaper model cost 14 times more in this case, on top of failing to finish the task. The example illustrates the study's main conclusion in a single instance: the cheaper model not only took more steps but spent tokens at every step — turning a $1 task into a $14 one.

The Cause: Thinking Tokens

The study's authors see the main cause of the phenomenon in the sharp variation in the amount of thinking tokens used by reasoning models. When answering a single prompt, different models' thinking-token consumption differs by up to 900%. Even with a low list price per token, the more a model "thinks" to shape its answer, the higher the final bill.

The same mechanism is visible in the Gemini example: the cheaper model took nearly 1,000 steps instead of 85 — spending far more thinking tokens to shape its answer. The pricier model, having "thought" less, completed the task faster and cheaper. Moreover, results are not stable even when the same prompt is resubmitted repeatedly: token consumption across repeated attempts varies by up to 9.7x. In other words, the same prompt can be "priced" differently each time. Together, these two factors fundamentally complicate accurate prediction of an individual prompt's cost — which is exactly the authors' conclusion.

Companies Respond

A Google spokesperson told WSJ that prompt-level variability is an inherent property of AI. According to the spokesperson, the company offers customers spending limits and flexible pricing plans.

"Prompt-level variability is an inherent property of AI," a Google spokesperson told WSJ. The spokesperson added that the company offers customers spending limits and flexible pricing plans.