| Model | Correct of 400 | 95% interval | Invalid | Per 1,000 | p50 | Confidence signal |
|---|---|---|---|---|---|---|
| Claude Opus 5.5 | 354 | 85.0%–91.3% | 0 | $4.03 | 3,018 ms | no |
| Gemini 3.8 Flash | 331 | 78.7%–86.1% | 0 | $1.04 | 2,774 ms | no |
| GPT-6 Sol | 330 | 78.5%–85.9% | 0 | $1.11 | 1,509 ms | no |
| gpt-oss-120b | 329 | 78.2%–85.7% | 2 | $0.0229 | 1,280 ms | yes |
| MiMo-V2.6-Flash | 326 | 77.4%–85.0% | 1 | $0.0299 | 2,790 ms | no |
| Kimi K2.6 | 323 | 76.6%–84.3% | 1 | $0.0847 | 719 ms | yes |
| GLM 5.3 Flash | 320 | 75.8%–83.6% | 1 | $0.0995 | 1,687 ms | yes |
| GPT-6 Luna | 320 | 75.8%–83.6% | 0 | $0.0556 | 902 ms | no |
| DeepSeek V4.1 Flash | 316 | 74.7%–82.7% | 0 | $0.047 | 1,021 ms | yes |
| MiniMax M3 | 316 | 74.7%–82.7% | 1 | $0.14 | 1,529 ms | yes |
| Qwen3.8 Flash | 316 | 74.7%–82.7% | 0 | $0.0466 | 1,041 ms | yes |
| gpt-oss-20b | 313 | 73.9%–82.0% | 1 | $0.014 | 1,355 ms | yes |
| Claude Sonnet 5 | 308 | 72.6%–80.9% | 0 | $1.89 | 2,375 ms | no |
| Gemini 3.5 Flash Lite | 305 | 71.8%–80.2% | 0 | $0.211 | 830 ms | no |
| Llama 3.3 70B Instruct | 305 | 71.8%–80.2% | 1 | $0.0676 | 946 ms | yes |
| Gemma 4 31B | 304 | 71.6%–79.9% | 1 | $0.0381 | 791 ms | yes |
| Claude Haiku 4.5 | 301 | 70.8%–79.2% | 2 | $0.691 | 986 ms | no |
| DeepSeek V4 Flash 0731 | 298 | 70.0%–78.5% | 1 | $0.0101 | 1,035 ms | yes |
| Llama 4 Maverick | 294 | 69.0%–77.6% | 1 | $0.101 | 1,154 ms | yes |
| Qwen3.7 Flash | 286 | 66.9%–75.7% | 6 | $0.0112 | 650 ms | yes |
| Qwen3.8 27B | 264 | 61.2%–70.5% | 82 | $0.16 | 1,634 ms | yes |
| Nemotron 3.5 Lightning | 255 | 58.9%–68.3% | 9 | $0.0245 | 934 ms | yes |
Could a cheap model replace the frontier model?
- GPT-6 Sol: 16 answers worse and 15 better on 400 fresh messages, 98.0% cheaper. not confirmed
- Gemini 3.8 Flash: 21 answers worse and 16 better on 400 fresh messages, 97.8% cheaper. not confirmed
On these messages the best cheap model got the same number right as GPT-6 Sol and Gemini 3.8 Flash, but on different messages. Our rule allows at most three answers made worse, so neither route shipped. Run twice, GPT-6 Sol changed only 7 of 400 answers, so the rule is reachable.
Limits: public dataset since 2020, so models may have seen it; some labels are ambiguous; single-label classification only.