LLM cost gateway · proof before routing

Cut the model bill.Keep every answer level.

Gimbix sits between your app and your AI providers. It moves a request to a cheaper model only when your own traffic proves the answer stays the same, then signs a receipt for every dollar saved.

−97.4%cost against the direct bill
142 vs 141correct, cheap route vs frontier
0answers made worse

Support-ticket benchmark, 148 fresh cases. Savings depend on the workload; our banking-intent test did not pass, and the ledger below says so.

How it holds level

Cheaper models only where your traffic proves they answer the same

Observe

Gimbix mirrors a share of your requests to candidate models and grades their answers against the model you use today. Nothing changes for your users.

Prove

A candidate must match your current model on one half of your traffic, then make at most three answers worse on the other half. The split is fixed by a salted hash, never by time of day.

Route

Only proven routes carry traffic. Approvals expire after seven days, and routing reverts on its own if quality drifts.

Receipt

Every saving is written to a hash-chained ledger signed with Ed25519. Your auditor verifies it offline with one public key, without trusting us.

Evidence ledger

What we measured, including what failed

Every row comes from a run whose rules were written down before the data existed. When a cheap route answers differently from your model, it does not ship, even when the total score is the same.

WorkloadRoute and testAnswersCostVerdict
Support triageLuna answers first, Sol takes low-confidence cases (selection, 400 held-out)391 vs 390−95.9%selected
Support triageSame route on 148 fresh cases142 vs 141, 0 worse−97.4%confirmed
Support triageThrough the real gateway1 worse, 2 better−97.8%confirmed
Support triageGuarded cloud run with an independent meter31 of 31 correct in both arms−95.2%pass
BANKING77 intentsCheapest route matching GPT-6 Sol on 400 fresh messages16 worse, 15 better−98.0%not confirmed
BANKING77 intentsCheapest route matching Gemini 3.8 Flash on 400 fresh messages21 worse, 16 better−97.8%not confirmed
BANKING77 intentsCheapest route matching Claude Opus 5.5 (selection)354 vs 354, 259 of 400 escalated−31.9%selected

GPT-6 Sol run twice against itself on the banking messages: 7 of 400 answers changed, so the three-regression rule is reachable and the failures above are real differences.

The confidence signal

Which models can answer first

A cheap model answers first only when it can say how sure it is. Gimbix reads that from token log probabilities. Of the 345 models in today's catalogue, 132 expose them through at least one host. Through OpenRouter, no OpenAI, Anthropic or Google model does, which is why Claude traffic needs a cross-provider first answer.

Browse the model price index or compare two models.

Example receipt

receipt      #1042
tenant       acme-support
window       2026-09-21 → 2026-09-28
route        luna → sol when confidence < 0.999
approval     cg-df5004f9…  expires 2026-10-05
requests     48,210   escalated 4.1%
baseline     $1,914.22
billed       $51.67
saved        $1,862.55
statement    sha256 9c1e…a07d
signature    ed25519 ok

Illustrative values. Real receipts carry counts, digests and signatures, never prompt or answer text.

See your own number before you switch anything

A pilot mirrors your traffic, proves routes on your requests and shows the saving your workload actually allows. Savings differ by workload; that is the point of measuring.

Request a pilot