Every agent team adds fallbacks eventually. Most of them are never exercised until the day they’re needed, and then they fail in new and interesting ways.
Fall back to something different
A fallback to the same provider in a different region protects you from a regional outage, not from a provider-wide one. Good chains mix providers and model sizes.
Test them every day
Send a small share of real traffic through each fallback on purpose. If the fallback breaks, you find out on a quiet Tuesday instead of during an outage.
ts
policy: { models: ["primary", "backup-a", "backup-b"], exercise: { share: 0.01 }, }
Know when you’re degraded
Tag every response with the model that produced it.
Alert when more than 5% of traffic runs on a fallback for more than 10 minutes.
Show degraded mode in your own UI when it makes sense.
Fallbacks are a feature, not a safety net. Treat them like one and they’ll be there when you need them.



