01

Product

How we cut token spend by 40%

Our own support agent was burning through tokens. Three changes to prompts, caching and routing cut the bill by 40% in two weeks.

Lina Park

Founding Engineer

8 min read

02

Dithered bar chart rising left to right with a dashed orange budget line and the tallest bar filled in orange

We run Sigil’s own support on an agent built with Sigil. In June, its token bill doubled in a month without any change in ticket volume. Here’s what we found and what we changed.

Find where the tokens go

Trace search made the first step easy. We grouped spans by agent step and sorted by tokens. Two steps accounted for 71% of spend: the system prompt, sent in full on every turn, and a documentation lookup that returned whole pages.

Change 1: cache the stable parts

Our system prompt and tool schemas are 6,000 tokens and almost never change. Moving them into a cached prefix cut input tokens per turn by more than half.

Change 2: return less from tools

The docs tool now returns the three most relevant sections instead of full pages. Answers got slightly better, because the model had less noise to read.

ts

docs.search(query, { limit: 3, maxTokens: 800, })

Change 3: route the easy turns

About half of support turns are simple follow-ups like “thanks” or “can you send that link again”. These now go to a small model.

Forty percent less spend, and the average answer got shorter and clearer.

None of these changes needed a new model or a rewrite. They needed visibility, and then a few hours of work.

03

04

Ship agents
that don’t
break.

Start free. See every run from day one.

RENDERING WORDMARK000%

SIGIL

Get it built

For you

Create a free website with Framer, the website builder loved by startups, designers and agencies.