Multi-model agents and routers
Most production agents aren’t “one model for every task.” A cheap model triages or retrieves, a stronger one reasons over the hard cases, or a router picks per request — mixing models is usually more cost-efficient than picking one model for everything, not less. So where does that fit?
Today: the whole flow is one configuration
Section titled “Today: the whole flow is one configuration”ParetoOps doesn’t need to know how a result was produced — only what it cost in total and whether
the task passed. That means a multi-model flow slots in exactly like a single-model one: give it
one configId (e.g. router-v2), and record, per task:
costUsd: the sum of every model call the flow made for that task (triage + reasoning + anything else).passed: whether the flow’s final answer was correct.taskId: so it can be paired against a baseline, same as any other config.
{ "configId": "router-v2", "configName": "Haiku-triage + Sonnet-reasoning", "taskId": "ticket-042", "model": "router-v2", "provider": "custom", "costUsd": 0.0043, "passed": true }Every feature already works on this: analyze places router-v2 on the frontier against
single-model configs on cost-per-successful-task; gate --baseline catches a regression if a
router change makes it worse or pricier; watch can early-stop a losing router variant the same
way. The comparison is honest precisely because it doesn’t care what’s inside the box — a
3-call router and a 1-call model are judged on the same two numbers everyone actually pays for.
One thing that doesn’t apply automatically: --reprice
Section titled “One thing that doesn’t apply automatically: --reprice”gate --reprice recomputes cost from token counts against the
pricing catalog, matched by model + provider. A composite flow’s model field (router-v2
above) has no catalog entry, so --reprice leaves its costUsd as originally recorded rather than
guessing at a blended rate. Two ways to still get that benefit for a router config:
- Keep recording the actual total spend per task (what the example above already does) — this is already correct and doesn’t need repricing.
- Or register your own blended rate for the virtual model name in
pareto.pricing.json(pareto-ops pricing:set --model router-v2 --provider custom --prompt <rate> --completion <rate>) if you want--repriceto recompute it from token totals instead.
What this doesn’t do yet: optimizing within a flow
Section titled “What this doesn’t do yet: optimizing within a flow”Comparing whole flows tells you that router-v2 beats always-sonnet. It doesn’t tell you
which step inside router-v2 is the one worth changing — e.g. “triage could use an even cheaper
model without losing accuracy.” That’s a per-task-segment routing analysis, not yet built (tracked
as a research-backed roadmap item: a routing simulator over FrugalGPT/RouteLLM/RouterBench-style
savings). Until then, the way to evaluate a routing change is the same as any other change: name
the new routing policy a new configId, run it, and let the regression gate and frontier tell you
whether it’s actually better.