LLM cost gateway · proof before routing
Cut the model bill.Keep every answer level.
Gimbix sits between your app and your AI providers. It moves a request to a cheaper model only when your own traffic proves the answer stays the same, then signs a receipt for every dollar saved.
Support-ticket benchmark, 148 fresh cases. Savings depend on the workload; our banking-intent test did not pass, and the ledger below says so.
How it holds level
Cheaper models only where your traffic proves they answer the same
Observe
Gimbix mirrors a share of your requests to candidate models and grades their answers against the model you use today. Nothing changes for your users.
Prove
A candidate must match your current model on one half of your traffic, then make at most three answers worse on the other half. The split is fixed by a salted hash, never by time of day.
Route
Only proven routes carry traffic. Approvals expire after seven days, and routing reverts on its own if quality drifts.
Receipt
Every saving is written to a hash-chained ledger signed with Ed25519. Your auditor verifies it offline with one public key, without trusting us.
Evidence ledger
What we measured, including what failed
Every row comes from a run whose rules were written down before the data existed. When a cheap route answers differently from your model, it does not ship, even when the total score is the same.
| Workload | Route and test | Answers | Cost | Verdict |
|---|---|---|---|---|
| Support triage | Luna answers first, Sol takes low-confidence cases (selection, 400 held-out) | 391 vs 390 | −95.9% | selected |
| Support triage | Same route on 148 fresh cases | 142 vs 141, 0 worse | −97.4% | confirmed |
| Support triage | Through the real gateway | 1 worse, 2 better | −97.8% | confirmed |
| Support triage | Guarded cloud run with an independent meter | 31 of 31 correct in both arms | −95.2% | pass |
| BANKING77 intents | Cheapest route matching GPT-6 Sol on 400 fresh messages | 16 worse, 15 better | −98.0% | not confirmed |
| BANKING77 intents | Cheapest route matching Gemini 3.8 Flash on 400 fresh messages | 21 worse, 16 better | −97.8% | not confirmed |
| BANKING77 intents | Cheapest route matching Claude Opus 5.5 (selection) | 354 vs 354, 259 of 400 escalated | −31.9% | selected |
GPT-6 Sol run twice against itself on the banking messages: 7 of 400 answers changed, so the three-regression rule is reachable and the failures above are real differences.
The confidence signal
Which models can answer first
A cheap model answers first only when it can say how sure it is. Gimbix reads that from token log probabilities. Of the 345 models in today's catalogue, 132 expose them through at least one host. Through OpenRouter, no OpenAI, Anthropic or Google model does, which is why Claude traffic needs a cross-provider first answer.
Example receipt
receipt #1042 tenant acme-support window 2026-09-21 → 2026-09-28 route luna → sol when confidence < 0.999 approval cg-df5004f9… expires 2026-10-05 requests 48,210 escalated 4.1% baseline $1,914.22 billed $51.67 saved $1,862.55 statement sha256 9c1e…a07d signature ed25519 ok
Illustrative values. Real receipts carry counts, digests and signatures, never prompt or answer text.
See your own number before you switch anything
A pilot mirrors your traffic, proves routes on your requests and shows the saving your workload actually allows. Savings differ by workload; that is the point of measuring.
Request a pilot