Home / Evidence / Support triage

Support triage

Synthetic support tickets graded against a fixed rubric. GPT-5.6 Luna answers first with token log probabilities; if its least confident token is below 0.999, GPT-5.6 Sol answers instead.

Selection, held-out cases400
Correct, route vs Sol391 vs 390
Escalated to Sol67
Below direct bill95.9%

Confirmation on fresh cases

Fresh cases148
Correct, route vs Sol142 vs 141
Answers made worse0
Below direct bill97.4%

Through the gateway and in a guarded cloud run

Through the real gateway the same cases came in 97.8% below the direct bill (1 worse, 2 better). In a guarded cloud run with an independent meter, 31 of 31 pairs were correct in both arms, 30 were served by Luna and 1 escalated, 95.2% below the same-run direct arm; the meter matched 30 of 30 served answers.

Limits: one synthetic workload from one generator, one provider, latency from one client in Dubai.