AI spend under
adult supervision
GraySmith is the AI spend governance platform for the whole company, not just the engineering team. We find the AI spend nobody approved, route what's left to the cheapest model that still passes, and reconcile the savings against your provider invoice.
- Read-only by default
- No prompt content stored
- EU data residency
- Results reconciled to your invoice
Month-to-date spend
€70,280
57.4% below baseline
Savings
€94,720
reconciled to invoice
Cache hit-rate
38.1%
+4.2 pts week on week
Quality delta
+0.2%
7-day A/B active
Demo data. customer-042, sample cockpit view.
Works with OpenAI, Anthropic, Google Vertex, AWS Bedrock, Azure OpenAI, Mistral, Groq, Together, Cohere, Fireworks, Replicate and self-hosted endpoints.
Your AI bill has a grey zone.
Somewhere between the spend you approved and the spend you'd forbid sits a third category: the spend nobody decided on. A team that bought a tool on a card. An agent still running after the project shipped. An API key belonging to someone who left in March. Two departments paying for the same model. Retries, re-prompts and context nobody counted.
It isn't fraud and it isn't negligence. It's what happens when spend grows faster than the process around it. And it doesn't show up in the tools built for production inference, because those tools only see what engineering pointed them at.
Optimisation is a feature. Governance is a mandate.
Discover
Every AI tool, key, agent and seat in the company, including the ones IT never provisioned. We reconcile provider accounts, card and expense data, SSO logs and network egress into one inventory, with an owner against every line.
Attribute
Every token mapped to a person, team, feature, customer and environment. Chargeback and showback, exported to your finance system or as CSV, reconciled monthly against the raw provider invoice, not an estimate.
Route
One endpoint in front of every provider. Each request scored and routed on task type, quality bar, latency budget and live price. Cheap models handle what they can. Hard requests escalate. Quality is measured, not assumed, and routing rolls back automatically if it drops.
Govern
Per-team and per-agent budgets with hard ceilings. Runaway-loop detection with automatic kill. Data-residency and provider policy enforced per request. Approval workflow for new tools. Alerts into Slack, PagerDuty, email or webhook.
Four weeks from scan to verified savings.
| Week | What happens | What you do |
|---|---|---|
| 0 | Free discovery scan, read-only. You get the inventory whether or not you continue. | One hour |
| 1 | Findings review. Kill the spend with no owner. | Two hours |
| 2 | Attribution live. Dashboards, chargeback, alerts. | Half a day of engineering time |
| 3 to 4 | Router deployed behind a feature flag, A/B tested for seven days minimum. | Review the results |
| Monthly | Reconciliation against the provider invoice. | Thirty minutes |
Six levers, deployed in order of risk.
Ranges are what these levers typically return on multi-provider workloads. Your number comes from the scan, and only verified savings, the ones visible on the provider invoice, count toward what you pay us.
- Unowned and orphaned spend10 to 25%
- Model routing30 to 50%
- Semantic caching20 to 40%
- Prompt and context compression15 to 30%
- Batch and async pricing10 to 50%
- Provider arbitrage and fallbacks20 to 35%
- Providers
- OpenAI, Anthropic, Google Vertex AI, AWS Bedrock, Azure OpenAI, Mistral, Groq, Together, Cohere, Fireworks, Replicate, OSS endpoints
- Finance
- NetSuite, QuickBooks, Xero, SAP, CSV, API
- Identity and spend
- Okta, Entra ID, Google Workspace, Pleo, Revolut Business, corporate card feeds
- Alerting
- Slack, Microsoft Teams, PagerDuty, email, webhook
- Telemetry
- OpenTelemetry ingest, one line of code
You can't optimise what you haven't found.
The scan is free, read-only, and takes an hour of your time. You keep the findings either way.