Which model is cheaper for your task?
Neither the smaller model nor the frontier model always wins. It depends on how hard the task is, how much context you send, how much the model thinks, what a failure costs downstream, and the success rate you need. Pick a scenario, select model presets, or customize your own numbers.
- Cost per call:
- Model succeeds, after retries:
- Needs the fallback:
- Cost per call:
- Model succeeds, after retries:
- Needs the fallback:
Live Pareto Plot
Visualizing cost vs. accuracy trade-off between the two modelsChange the scenario, model presets, context size or failure cost, and the answer can flip.
Stop estimating. Let your evals measure the frontier.
The calculator above uses estimated success rates and token counts. In production, prompt rewrites move results 11–58× more than rerunning them, and token spend varies up to 30× across tasks.
ParetoOps replaces manual guesswork with deterministic CI analysis across all your prompts and models:
- Reads your real eval exports: Directly ingests results from Promptfoo, LangSmith, Braintrust, DeepEval, CSV, and JSON.
- Discovers the true Pareto frontier: Evaluates dozens of models and prompt variants at once, identifying dominated options and statistical sweet spots with 95% Wilson confidence intervals.
- Guards your budget in CI: Exits 1 if a pull request regresses accuracy or increases cost per success, posting an explanatory diff right on the PR.
$ npx pareto-ops analyze eval_benchmark.json --min-acc 0.90 Config / Model Success % 95% Wilson Cost/Success Status ---------------------------------------------------------------------- Gemini 2.5 Flash 67.0% 57% min $0.0022 ✔ OPTIMAL Claude Haiku 4.5 92.0% 85% min $0.0035 ✔ OPTIMAL Claude Sonnet 5 94.0% 88% min $0.0233 ✔ OPTIMAL Claude Opus 4.8 80.0% 71% min $0.0327 ✖ DOMINATED 🌟 Recommended Sweet Spot: Claude Haiku 4.5 at $0.0035/success Meets 90.0% accuracy bar; 85% cheaper than Sonnet 5 $ npx pareto-ops gate eval_benchmark.json \ --config claude-haiku-4-5 --min-acc 0.90 ✅ CI Gate PASSED: 0 regressions detected.
Runs locally or in CI — no telemetry, your data stays with you.
What moves the answer
Each of these can make the smaller model the right choice, or the frontier model.
Task complexity
On easy tasks, models score alike and the cheaper one wins. On hard ones, success rates drift apart, and failures dominate the cost. Different kinds of models turn out to be the most economical in different domains (Erol et al., 2025).
Prompt and context size
Long instructions, documents and history are sent on every call and multiply the price gap between models. In agent workloads, input tokens drove most of the cost (Bai et al., 2026).
How much the model thinks
Reasoning tokens are billed as output and vary widely: between models, and between runs of the same query. That is how a model with a lower list price can end up costing more (Chen, Zhang et al., 2026).
What a failure costs you
If a miss is dropped or retried cheaply, a less reliable model can be fine. If it goes to a support agent or an engineer, reliability is worth paying for.
The success rate you need
A model below your bar is out, however cheap. Above it, extra accuracy may be worth little: accuracy often plateaus while spend keeps rising (Bai et al., 2026).
Retries, volume and latency
Retries lift success but add calls and wait time, and a hard task that failed once often fails again, so treat retry gains as optimistic. At volume, fractions of a cent become real money.
How it's calculated
With first-try success rate p and up to n attempts, a task fails every attempt with probability (1 − p)n.
cost per call = context tokens × input price + output tokens × output pricemodel success = 1 − (1 − p)ⁿexpected calls per task = model success ÷ p- Failures completed by a fallback:
cost per successful task = cost per call × expected calls + (1 − p)ⁿ × cost per failure - Failures accepted:
cost per successful task = cost per call ÷ p
It assumes attempts succeed independently, which flatters retries. And it needs a success rate you can trust: measured on many tasks, on your prompts. A small eval can be off by more than the gap between two models.
ParetoOps does this from your real eval results, for every model and prompt you test at once, and shows which ones are worth keeping for each job. See how →
Stop estimating. Measure it on your tasks.
ParetoOps Free computes this for every configuration in your eval results. No sign-up.