The idea
What is cost per successful task?
It is the total you spent on a set of tasks divided by the number of tasks that actually succeeded. A failed call still costs money, so a model that is cheap per call but fails often can cost more per useful result than a pricier, more reliable one. Researchers call the same quantity cost-of-pass.
Why not just compare price per token?
Because price per token says nothing about how many tokens a model spends on your task, or how often it gets the task right. Published research has found that the model with the lower list price can cost more in total, sometimes many times more, mostly because of how many reasoning tokens it uses.
Should I always use the most capable model?
Not necessarily. Research that priced each correct answer found that different kinds of models are the most economical for different tasks, and that sending each request to the cheapest model able to handle it can match the best single model at a fraction of the cost. On hard tasks where failures are expensive, the most capable model can still be the cheapest per success. The right choice depends on task complexity, prompt and context size, how much the model reasons, what a failure costs you and the success rate you need, so it has to be measured on your own tasks.
What is a Pareto frontier, and what does dominated mean?
A configuration is dominated when another one is at least as accurate and no more expensive per successful task, and strictly better on one of the two. The configurations nothing dominates form the Pareto frontier. Each frontier point is a reasonable choice; everything off it costs more for the same or worse results.
How many tasks do I need for a trustworthy result?
More than most people expect. Success rates measured on a few dozen tasks are noisy, and the same task can cost very different amounts from run to run. ParetoOps shows a 95% lower bound next to each success rate so you can see how much to trust it. Repeating each task a few times also helps.
Using ParetoOps
Does ParetoOps run my evals or call model APIs?
No. It reads the results your eval tool already produces and does the analysis on top. It never calls a model provider.
Which eval tools does it work with?
Promptfoo, LangSmith, Braintrust and DeepEval exports, plus CSV and a simple JSON format. Each trial needs a configuration name, its cost and whether it passed. Latency is optional.
Does my data leave my machine?
No. ParetoOps runs on your laptop or CI runner, sends no telemetry, and checks license keys offline.
Does it account for prompt caching and reasoning tokens?
Yes. Its pricing catalog includes cache-read, cache-write and reasoning-token rates for the major providers, alongside the standard input and output prices.
Can I use it in CI for free?
Yes. The free gate command fails a CI job when a configuration drops below your accuracy threshold or goes over your cost-per-success budget. It works in any CI system that can run a command.
Is there a Python SDK?
Yes. Install it with pip install pareto-ops to compute the frontier and recommendations from Python.
Plans
What does Pro add?
The checks a team needs on every pull request: a CI gate that comments on the pull request (a GitHub Actions step, or any CI), a regression gate that compares each change with your last known good run task by task, baselines stored on a git branch, early stopping for eval runs that are already failing, and a 3D frontier that includes P95 latency.
When is Pro available?
Pro and Enterprise are coming soon. Join the interest list to hear when they launch. People on the list may be offered a trial key.
How is Pro licensed?
Per GitHub organization, with a license key you store as a secret. The key is verified offline: nothing calls home.
Do I need an account to use Free?
No. Free needs no key and no sign-up.