Large models are great, and expensive. In most production agents we look at, fewer than a third of requests actually need the largest model. The rest are classification, extraction or short replies that a small model handles just as well.
Start with a baseline
Before routing anything, measure. Replay a week of real traffic through both models and score the answers. Sigil’s shadow routes do this without touching production.
A policy that works
The simplest policy that works well is: try the small model first, and escalate when confidence is low or a tool call fails.
ts
policy: { models: ["small", "large"], escalateWhen: ["low_confidence", "tool_error"], }
What to watch
Escalation rate: if more than 40% of requests escalate, the small model is the wrong fit.
Latency: two calls are slower than one, so track p95 end to end.
Quality: keep scoring a sample of routed answers every week.
Results we see
Teams that adopt this pattern typically move 60–75% of traffic to small models and cut model spend by a third, with no measurable change in answer quality.



