We run Sigil’s own support on an agent built with Sigil. In June, its token bill doubled in a month without any change in ticket volume. Here’s what we found and what we changed.
Find where the tokens go
Trace search made the first step easy. We grouped spans by agent step and sorted by tokens. Two steps accounted for 71% of spend: the system prompt, sent in full on every turn, and a documentation lookup that returned whole pages.
Change 1: cache the stable parts
Our system prompt and tool schemas are 6,000 tokens and almost never change. Moving them into a cached prefix cut input tokens per turn by more than half.
Change 2: return less from tools
The docs tool now returns the three most relevant sections instead of full pages. Answers got slightly better, because the model had less noise to read.
ts
docs.search(query, { limit: 3, maxTokens: 800, })
Change 3: route the easy turns
About half of support turns are simple follow-ups like “thanks” or “can you send that link again”. These now go to a small model.
Forty percent less spend, and the average answer got shorter and clearer.
None of these changes needed a new model or a rewrite. They needed visibility, and then a few hours of work.



