SRE Bench Update: The Models Stopped Failing

SRE Bench measures one thing: give an AI model a live incident in a real broken environment, through the same tools an on-call engineer uses, and check whether it names the true root cause. A pinned judge grades every investigation against written ground truth. Naming the broken component is not enough.
In August that bar was hard to clear. Our trap incidents, where the obvious first answer is wrong, split the field cleanly. One scenario alone defeated seven models. The first champion lost its crown within 48 hours when we added two harder tasks.
October looks different.
What Changed
We benchmarked the new OpenAI generation this week. GPT-6 Sol, GPT-6 Luna, and GPT-6.1 Sol all scored 100 percent on the first run. No partial credit, no misses, traps included.
The detail worth staring at is Luna. Its predecessor, GPT-5.6 Luna, sat in last place on the August board. GPT-6 Luna is perfect, and it found each root cause for about $0.003. That is the cheapest correct answer we have ever recorded, roughly 280 times cheaper per root cause than Claude Fable 5, which is also perfect.
The same week, Claude Opus 5.5 went perfect at one sixth the cost of Claude Opus 5. A month earlier, Gemini 3.7 Flash went from zero to first place once we routed around a broken API translation layer. The pattern is consistent: every frontier release clears our bar, and each one clears it cheaper and faster than the last.
What It Means for Operators
Three practical conclusions from the current board:
- Accuracy on routine incidents is now a commodity. Crash loops, broken configs, dead services, resource exhaustion: a dozen models from seven vendors diagnose these correctly. If you pay flagship prices for routine triage, you are paying for a margin that no longer exists.
- Cost and speed are the real ranking. Perfect scorers span $0.003 to $0.85 per root cause and 49 to 126 seconds per investigation. That is a 280x price spread for the same answer.
- Reliability is the next frontier. A single clean run is not the same as 5 clean runs out of 5. On-call does not reward the model that is right once.
What It Means for SRE Bench
A benchmark that everyone passes has stopped measuring. We are not going to pretend otherwise, and we are not going to quietly retire the hard results that came before.
The third edition is in design now: deeper multi-hop incidents in the style of our most discriminating scenario, where the visible failure is two causal layers away from the true cause, plus repeated runs so every score carries a reliability figure instead of a lucky point estimate. Models that merely pass once will separate from models you can page at 3 AM.
Conclusion
The models stopped failing our incidents, which is good news for operators and a deadline for benchmark authors. The full table, methodology, and per-model details are at akmatori.com/srebench.
Akmatori helps SRE teams connect alerts, metrics, logs, traces, and runbooks into governed incident workflows. Powered by Gcore, Akmatori gives operators the control they need when production behavior needs a fast explanation.
