Skip to content

Multi-model agents and routers

Most production agents aren’t “one model for every task.” A cheap model triages or retrieves, a stronger one reasons over the hard cases, or a router picks per request — mixing models is usually more cost-efficient than picking one model for everything, not less. So where does that fit?

Today: the whole flow is one configuration

Section titled “Today: the whole flow is one configuration”

ParetoOps doesn’t need to know how a result was produced — only what it cost in total and whether the task passed. That means a multi-model flow slots in exactly like a single-model one: give it one configId (e.g. router-v2), and record, per task:

  • costUsd: the sum of every model call the flow made for that task (triage + reasoning + anything else).
  • passed: whether the flow’s final answer was correct.
  • taskId: so it can be paired against a baseline, same as any other config.
{ "configId": "router-v2", "configName": "Haiku-triage + Sonnet-reasoning", "taskId": "ticket-042",
"model": "router-v2", "provider": "custom", "costUsd": 0.0043, "passed": true }

Every feature already works on this: analyze places router-v2 on the frontier against single-model configs on cost-per-successful-task; gate --baseline catches a regression if a router change makes it worse or pricier; watch can early-stop a losing router variant the same way. The comparison is honest precisely because it doesn’t care what’s inside the box — a 3-call router and a 1-call model are judged on the same two numbers everyone actually pays for.

One thing that doesn’t apply automatically: --reprice

Section titled “One thing that doesn’t apply automatically: --reprice”

gate --reprice recomputes cost from token counts against the pricing catalog, matched by model + provider. A composite flow’s model field (router-v2 above) has no catalog entry, so --reprice leaves its costUsd as originally recorded rather than guessing at a blended rate. Two ways to still get that benefit for a router config:

  • Keep recording the actual total spend per task (what the example above already does) — this is already correct and doesn’t need repricing.
  • Or register your own blended rate for the virtual model name in pareto.pricing.json (pareto-ops pricing:set --model router-v2 --provider custom --prompt <rate> --completion <rate>) if you want --reprice to recompute it from token totals instead.

What this doesn’t do yet: optimizing within a flow

Section titled “What this doesn’t do yet: optimizing within a flow”

Comparing whole flows tells you that router-v2 beats always-sonnet. It doesn’t tell you which step inside router-v2 is the one worth changing — e.g. “triage could use an even cheaper model without losing accuracy.” That’s a per-task-segment routing analysis, not yet built (tracked as a research-backed roadmap item: a routing simulator over FrugalGPT/RouteLLM/RouterBench-style savings). Until then, the way to evaluate a routing change is the same as any other change: name the new routing policy a new configId, run it, and let the regression gate and frontier tell you whether it’s actually better.