Pro

What ParetoOps Pro posts on a pull request

The Pro CI gate (pareto-ops ci) writes one report to the job summary and keeps one comment up to date on the pull request. Both reports below are its real output on the sample data shipped with ParetoOps, rendered the way GitHub shows them. No invented numbers, nothing redacted.

A pull request that passes

Four model configurations on a 100-task benchmark, with a 90% accuracy bar. Claude Haiku 4.5 clears the bar at 85% less cost per success than Claude Sonnet 5, so the report recommends it; the legacy Opus setup is dominated.

ParetoOpscommented on this pull requestโœ… PASSED

๐Ÿ“Š ParetoOps Gate Report โ€” โœ… PASSED

๐ŸŒŸ Recommended Sweet Spot

Claude Haiku 4.5 (claude-haiku-4-5) delivers optimal cost efficiency at $0.0035 per successful task with 92.0% task accuracy.

๐Ÿ“‰ Pareto Frontier

Cost per successful task vs success rateAccurate but costlySweet spotCheap but unreliableAvoidCheaper per successCostlier per successLess accurateMore accurateGemini 2.5 FlashClaude Haiku 4.5Claude Sonnet 5Claude Opus 4.8 legacy ยท dominated

๐Ÿ“ˆ Model Efficiency Matrix

ConfigurationTrialsSuccess RateAvg CostCost / SuccessStatus
Gemini 2.5 Flash10067.0%$0.0015$0.0022๐ŸŸข Optimal
Claude Haiku 4.510092.0%$0.0032$0.0035๐ŸŸข Optimal
Claude Sonnet 510094.0%$0.0219$0.0233๐ŸŸข Optimal
Claude Opus 4.8 (legacy)10080.0%$0.0261$0.0327๐Ÿ”ด Dominated

๐Ÿ’ก Optimization Actions

  • ๐ŸŒŸ Sweet Spot: Claude Haiku 4.5 delivers the best cost efficiency at 92% success rate and $0.0035/successful task โ€” 85% cheaper than Claude Sonnet 5.
  • ๐Ÿ›‘ Eliminate: Claude Opus 4.8 (legacy) is mathematically dominated by [Claude Haiku 4.5, Claude Sonnet 5]. You are paying more per successful task for lower or identical reliability.
  • ๐Ÿ“‰ Cost Downgrade: Switching to Claude Haiku 4.5 saves ~85% on cost-per-successful-task with only a 2.0% delta in task success rate.
  • ๐Ÿ“ˆ Reliability Upgrade: Claude Opus 4.8 (legacy) achieves 80% success (below 90% SLA). Upgrading to Claude Haiku 4.5 gains +12.0% success rate and meets your SLA threshold.

Generated by ParetoOps โ€” cost per successful task, gated in CI.

Frontier report from examples/eval_benchmark.json (pareto-ops ci --config claude-haiku-4-5 --min-acc 0.9). GitHub draws the chart from the report's Mermaid block.

A regression caught before merge

The change still clears its 50% accuracy bar, so a threshold check alone would pass it. Compared task by task with the saved baseline, it is significantly less accurate and more expensive per success, so the gate fails the pull request.

ParetoOpscommented on this pull requestโŒ FAILED

๐Ÿ“Š ParetoOps Gate Report โ€” โŒ FAILED

โŒ Gate Violations

  • Regression vs baseline sonnet-5@v1: Accuracy is significantly lower: -15.0pp (CI -22.5pp to -7.5pp), beyond the 2.0pp margin. Cost per successful task is significantly higher: ร—1.42 (CI ร—1.24โ€“ร—1.64), beyond the +10% margin.

๐Ÿ” Change vs Baseline (200 shared tasks) โ€” WORSE

MetricBaselineThis changeChange (95% CI)Verdict
Accuracy81.5%66.5%-15.0pp (-22.5pp to -7.5pp)WORSE
Cost / Success$0.0155$0.0219ร—1.42 (ร—1.24โ€“ร—1.64)WORSE

๐Ÿ“ˆ Model Efficiency Matrix

ConfigurationTrialsSuccess RateAvg CostCost / SuccessStatus
Claude Sonnet 5 (prompt v1)20066.5%$0.0146$0.0219๐ŸŸข Optimal

Generated by ParetoOps โ€” cost per successful task, gated in CI.

Regression report from examples/ci/candidate-regressed.json against examples/ci/baseline.json (200 shared tasks).