01

Product

Routing between small and large models

Most requests don’t need your biggest model. A simple routing policy can send them to a small one without users noticing.

Aiko Tanaka

Product Lead

6 min read

02

A large dithered grey square connected by thin lines to a small orange dithered square, one line highlighted in orange

Large models are great, and expensive. In most production agents we look at, fewer than a third of requests actually need the largest model. The rest are classification, extraction or short replies that a small model handles just as well.

Start with a baseline

Before routing anything, measure. Replay a week of real traffic through both models and score the answers. Sigil’s shadow routes do this without touching production.

A policy that works

The simplest policy that works well is: try the small model first, and escalate when confidence is low or a tool call fails.

ts

policy: { models: ["small", "large"], escalateWhen: ["low_confidence", "tool_error"], }

What to watch

  • Escalation rate: if more than 40% of requests escalate, the small model is the wrong fit.

  • Latency: two calls are slower than one, so track p95 end to end.

  • Quality: keep scoring a sample of routed answers every week.

Results we see

Teams that adopt this pattern typically move 60–75% of traffic to small models and cut model spend by a third, with no measurable change in answer quality.

03

04

Ship agents
that don’t
break.

Start free. See every run from day one.

RENDERING WORDMARK000%

SIGIL

Get it built

For you

Create a free website with Framer, the website builder loved by startups, designers and agencies.