Six levers. Deployed in order of risk.
We start with the savings that have no downside and only touch quality when the A/B test says we can.
| Lever | Typical reduction | What it does |
|---|---|---|
| Kill unowned spend | 10 to 25% | Removes spend with no owner and no purpose |
| Model routing | 30 to 50% | Matches each request to the cheapest passing model |
| Semantic caching | 20 to 40% | Serves near-identical requests from cache |
| Prompt and context compression | 15 to 30% | Trims context that adds no measurable quality |
| Batch and async pricing | 10 to 50% | Moves latency-tolerant work to discounted tiers |
| Provider arbitrage and fallbacks | 20 to 35% | Shifts traffic on live price and availability |
Nothing ships without a test.
Every change runs behind a feature flag and is A/B tested for a minimum of seven days against the incumbent. We measure cost, latency and quality. If quality regresses beyond your threshold, the change rolls back automatically and we report it, including when the test says a lever doesn't work for you.
Savings are reconciled, not estimated.
At the end of each billing cycle we compare the provider's raw invoice against your pre-GraySmith baseline. That difference is the only number we count, the only number we report, and the only number you're billed a share of.