- Across 8 reasoning models and 12 tasks, the model with the lower list price cost more in total in 32% of model pairs, by as much as 28 times.
- In one pair, a model listed 80% cheaper than the other ended up 38% more expensive across the tasks.
- The main driver was how many thinking tokens each model spent. Running the same query again could change its thinking-token count by up to 9.7 times.
What it means in practice: Price per token does not predict what a task will cost you. The only reliable number is measured spend on your own tasks.
- Runs of the same agentic coding task varied by up to 30 times in total tokens.
- Input tokens, not output tokens, drove most of the cost.
- Accuracy often peaked at a moderate level of spend and then flattened out as spend kept rising.
- Models were poor at predicting their own token usage, with correlations of 0.39 at best.
What it means in practice: One run tells you little about cost: measure across repeated trials. And past a certain point, paying more stops buying accuracy.
- Defines cost-of-pass: the expected money spent to get one correct solution. It is the same idea ParetoOps calls cost per successful task.
- Different kinds of models were the most economical choice in different domains: lightweight models for basic quantitative tasks, large models for knowledge-heavy tasks, reasoning models for complex quantitative ones.
- For complex quantitative tasks, the cost of a correct answer roughly halved every few months over the year studied.
What it means in practice: The best model depends on your workload, and the answer goes stale quickly. It is worth re-measuring whenever models or prices change.
- Benchmarks that ranked agents on accuracy alone encouraged agents that were needlessly complex and expensive.
- Optimizing cost and accuracy together cut cost substantially while keeping accuracy.
What it means in practice: Judge configurations on cost and accuracy at the same time, and look for the ones nothing else beats on both.
- Sending each query to the cheapest model likely to handle it, and escalating only when needed, matched the best single model of the time at up to 98% lower cost.
- Spending the same budget that way instead raised accuracy by 4% over that best single model.
What it means in practice: The strongest model is not automatically the right one for every request. A mix chosen per task can beat any single model on cost, accuracy or both.
- Routing each query to either a stronger or a weaker model cut costs by more than half in some cases, without lowering response quality.
- The routers kept working when the underlying models were swapped.
What it means in practice: Many requests do not need the most capable model. Knowing which ones do is where the savings are.
- Treats an eval as an experiment on a sample of possible questions, so every score comes with uncertainty.
- Gives methods for measuring the difference between two models, and for planning evals with enough statistical power.
What it means in practice: Comparing two eval scores by eye is not enough to say a change helped or hurt. The difference needs an error bar.
- Rewording a prompt without changing its meaning moved results 11 to 58 times more than simply rerunning the same prompt (median paired standard deviation).
What it means in practice: Small prompt edits can swing results far more than run-to-run noise suggests. Regressions hide easily unless you compare task by task.
- Keeping only tasks with a middling historical pass rate (30 to 70%) cut the number of eval tasks by 44 to 70% while keeping model rankings largely intact.
What it means in practice: A large part of an eval budget can go to tasks that do not tell models apart.