GraySmith.io

AI spend under
adult supervision

GraySmith is the AI spend governance platform for the whole company, not just the engineering team. We find the AI spend nobody approved, route what's left to the cheapest model that still passes, and reconcile the savings against your provider invoice.

  • Read-only by default
  • No prompt content stored
  • EU data residency
  • Results reconciled to your invoice

Works with OpenAI, Anthropic, Google Vertex, AWS Bedrock, Azure OpenAI, Mistral, Groq, Together, Cohere, Fireworks, Replicate and self-hosted endpoints.

Your AI bill has a grey zone.

Somewhere between the spend you approved and the spend you'd forbid sits a third category: the spend nobody decided on. A team that bought a tool on a card. An agent still running after the project shipped. An API key belonging to someone who left in March. Two departments paying for the same model. Retries, re-prompts and context nobody counted.

It isn't fraud and it isn't negligence. It's what happens when spend grows faster than the process around it. And it doesn't show up in the tools built for production inference, because those tools only see what engineering pointed them at.

Optimisation is a feature. Governance is a mandate.

Four weeks from scan to verified savings.

WeekWhat happensWhat you do
0Free discovery scan, read-only. You get the inventory whether or not you continue.One hour
1Findings review. Kill the spend with no owner.Two hours
2Attribution live. Dashboards, chargeback, alerts.Half a day of engineering time
3 to 4Router deployed behind a feature flag, A/B tested for seven days minimum.Review the results
MonthlyReconciliation against the provider invoice.Thirty minutes

Six levers, deployed in order of risk.

Ranges are what these levers typically return on multi-provider workloads. Your number comes from the scan, and only verified savings, the ones visible on the provider invoice, count toward what you pay us.

  • Unowned and orphaned spend10 to 25%
  • Model routing30 to 50%
  • Semantic caching20 to 40%
  • Prompt and context compression15 to 30%
  • Batch and async pricing10 to 50%
  • Provider arbitrage and fallbacks20 to 35%
Providers
OpenAI, Anthropic, Google Vertex AI, AWS Bedrock, Azure OpenAI, Mistral, Groq, Together, Cohere, Fireworks, Replicate, OSS endpoints
Finance
NetSuite, QuickBooks, Xero, SAP, CSV, API
Identity and spend
Okta, Entra ID, Google Workspace, Pleo, Revolut Business, corporate card feeds
Alerting
Slack, Microsoft Teams, PagerDuty, email, webhook
Telemetry
OpenTelemetry ingest, one line of code

You can't optimise what you haven't found.

The scan is free, read-only, and takes an hour of your time. You keep the findings either way.