Skip to content

Your first analysis

ParetoOps reads results from Promptfoo, LangSmith, Braintrust, DeepEval, CSV, or its own JSON format. In the JSON format, each trial is one model call on one task:

[
{
"configId": "gemini-2.5-flash",
"configName": "Gemini 2.5 Flash",
"model": "gemini-2.5-flash",
"provider": "google",
"costUsd": 0.0018,
"passed": true,
"durationMs": 420
}
]
  • configId: the configuration being compared (a model, or a model plus prompt version).
  • costUsd: what the call cost.
  • passed: whether the task succeeded.
  • durationMs (optional): latency.

The format is detected automatically. To force it, pass --format promptfoo (or langsmith, braintrust, deepeval, csv, pareto).

Terminal window
npx pareto-ops analyze eval_benchmark.json --min-acc 0.9
Config / Model Success % 95% Wilson Avg Cost Cost/Success Status
--------------------------------------------------------------------------------
Gemini 2.5 Flash 67.0% 57% min $0.0015 $0.0022 ✔ OPTIMAL
Claude Haiku 4.5 92.0% 85% min $0.0032 $0.0035 ✔ OPTIMAL
Claude Sonnet 5 94.0% 88% min $0.0219 $0.0233 ✔ OPTIMAL
Claude Opus 4.8 (legac 80.0% 71% min $0.0261 $0.0327 ✖ DOMINATED
🌟 Recommended Sweet Spot: Claude Haiku 4.5 at $0.0035/successful task (92.0% accuracy)
  • Cost/Success is total spend divided by passed tasks: the number to compare.
  • 95% Wilson is a lower bound on the success rate. With few trials, the real rate could be this low. Run more tasks to tighten it.
  • OPTIMAL configurations are on the Pareto frontier: no other configuration is both cheaper per success and at least as accurate.
  • DOMINATED configurations are beaten on both axes by something else, so you’re paying more for the same or worse results. Here, the legacy Claude Opus 4.8 setup costs about nine times as much per success as Claude Haiku 4.5 and is less accurate.
  • Low sample warnings appear when a configuration has fewer than 10 trials: its 95% lower bound is then too loose to trust a recommendation. Run at least 30 trials per configuration.
  • Sweet spot is the cheapest frontier configuration that meets your accuracy bar, set with --min-acc (default 0.8). Here Claude Haiku 4.5 meets a 90% bar at 85% less cost per success than Claude Sonnet 5, which is only 2 points more accurate, so ParetoOps also suggests a downgrade.

Add --json to get the full analysis as JSON.

Add a CI gate.