Skip to main content
08.10.2026

Claude Haiku 5.5: Aces the Traps, Flunks the Gimme

head-image

SRE Bench gives each model a live incident in a real broken environment: a real alert, real tools, a fresh install of Akmatori with no accumulated memory. A pinned judge grades every investigation against written ground truth. Naming the broken component is not enough. The model must find the cause behind it.

This week we added Claude Haiku 5.5, Anthropic's budget tier. The results are worth a closer look than the score alone suggests.

The Numbers

  • Root causes found: 92 percent. Eleven clean identifications, one partial.
  • Cost per root cause: about $0.003. That ties GPT-6 Luna for the cheapest correct answer on the board, 26 times cheaper than Gemini 3.7 Flash.
  • Median investigation: 37 seconds. The fastest Anthropic model we have measured, and second-fastest overall across 25 entries.

For context, Claude Opus 5.5 is perfect on the same tasks at roughly 28 times the cost per answer and twice the wall-clock time.

It Passed the Hard Part

Our scenario suite includes trap incidents where the first plausible answer is wrong. A 502 where the web server is healthy and the true cause is a corrupted config value two layers down. A service with stale manually-managed endpoints while a healthy backend sits unselected. A decoy crash-loop planted next to the real problem.

These traps are what separate the field. One of them has defeated seven models. Haiku 5.5 cleared every one of them, including the disk-full red herring that stopped its predecessor, Claude Haiku 4.5.

It Failed the Easy Part

The partial came on the simplest scenario we run: a stopped web server. Nothing listens on port 80. The nginx service is inactive. One command settles it.

Haiku 5.5 correctly observed that nothing was listening and that no nginx process existed. Then, instead of checking the service state, it built a theory about long-decommissioned Docker containers and a stale monitoring probe. It never ran systemctl status nginx. All but one of the other models on the board ran it and closed the case.

Every operator has seen this failure in a human: the investigator who theorizes past the boring check. It is instructive to see it in a model that had just finished out-reasoning most of the field on genuinely hard incidents.

What Operators Should Take From This

  • The budget tier is no longer a toy. At $0.003 per root cause with 37-second investigations, Haiku 5.5 and GPT-6 Luna make routine triage effectively free.
  • At identical prices, the differentiator is reliability. Luna is perfect but takes 91 seconds. Haiku is 2.5 times faster but skipped a basic check once. For a paging workflow, that once matters.
  • Single runs have limits. Was the miss character or dice? One run cannot say. That is exactly why the next edition of SRE Bench adds repeated runs, so every score carries a reliability figure instead of a point estimate.

Conclusion

Haiku 5.5 is a remarkable result sold short by its own scorecard: trap-solving competence at commodity prices, undone by one skipped command on the easiest task we have. The full table, methodology, and every archived investigation are at akmatori.com/srebench.

Akmatori helps SRE teams connect alerts, metrics, logs, traces, and runbooks into governed incident workflows. Powered by Gcore, Akmatori gives operators the control they need when production behavior needs a fast explanation.

Automate incident response and prevent on-call burnout with AI-driven agents!