The AI SRE bake-off

Don't take our word for it.
Or anyone's.

AI SRE is the loudest category in infrastructure, and every demo you will see is rehearsed — including ours. Published benchmarks put frontier AI agents at 39–73% diagnosis success depending on the scenario. So the only evaluation that matters runs on your incidents, not a vendor's demo environment. Here is the rubric we ask prospects to score us on. Use it on everyone. Including us.

Seven scenarios every vendor should survive

Build each one from a real incident in your own history. If a product only shines on the easy ones, you want to know before the contract, not after.

01

Deploy regression

A failure storm starts minutes after a release. Somewhere in that deploy is the change that caused it.

Passing looks like: Names the specific pull request — not just "a recent deploy" — and the recommended action is the revert.

02

Shared-infra blast radius

One piece of shared infrastructure fails and every service that depends on it starts alerting at once.

Passing looks like: One incident with one cause, not a page per service. Attribution lands on the platform and the storm is deduplicated.

03

One symptom, many causes

The same timeout fires on three different days for three different upstream reasons.

Passing looks like: Distinguishes the three causes instead of pasting the same diagnosis three times.

04

The wrapped error

The log says unhandled-exception. The truth is two layers further down.

Passing looks like: Digs past the generic wrapper to the real cause instead of ranking the loudest label first.

05

Not your fault

The failure is on the customer side — expired credentials, a missing file, bad input.

Passing looks like: Attributes it to the customer, does not page the platform team, and proposes the fix where the fix belongs.

06

It healed itself

A transient blip self-recovers before any human could have acted on it.

Passing looks like: Stays silent or auto-closes. Not crying wolf is a scored skill, not a lucky outcome.

07

Never seen before

A novel failure with no precedent in any runbook, memory or past incident.

Passing looks like: A reasoned hypothesis with calibrated confidence — and the honesty to escalate to a human rather than bluff.

Ten dimensions, scored blind

Score every vendor on every scenario, by an engineer who knows what actually happened.

  • 01Root-cause correctnessNone, partial or full — scored against what your engineers concluded at the time.
  • 02AttributionPlatform fault or customer fault. Getting this wrong pages the wrong people.
  • 03Recommended actionRevert, retry, escalate, silence: was the proposed next step the right one?
  • 04Escalate vs. silenceDid it wake a human exactly when it should have — and only then?
  • 05Alert-storm deduplicationA hundred alerts, one cause: did it produce one incident or a hundred?
  • 06Confidence calibrationWhen it said 90%, was it right 90% of the time?
  • 07Evidence and citation qualityEvery claim traceable to a log line, a metric or a commit that a reviewer can check.
  • 08Time to root causeWall-clock from alert to a correct, usable diagnosis.
  • 09Cost per investigationWhat the run actually cost — visible during the pilot, not estimated after the invoice.
  • 10Hallucination countAnything invented — a service, a metric, a log line — is disqualifying, not a deduction.

The whole field, one sheet

Every vendor that shows up in AI SRE shortlists, on the same rows — including the rows we lose. A table where the host wins everything is marketing; this one you can check.

CapabilityWHAWITincident.ioPagerDutyRootlyDatadog BitsResolve.aiTraversalClericNeuBird
AI root-cause investigation with cited evidence
Incident management system of record: lifecycle, timeline, postmortems
Native on-call: schedules, rotations, escalation policies
Reads your existing observability stack read-only, no migration
AI agents with scheduled shifts on the on-call roster
Coding agent ships the fix as a pull request
Infrastructure remediation: rollback, scale, restart
WhatsApp as a paging channel
Public pricing, or AI spend itemized inside the product

✓ shipped  ·  ◐ partial, opt-in, preview or limited scope  ·  — not present in the vendor's public documentation. Cells reflect each vendor's public documentation as of August 2026. All product names are trademarks of their respective owners; none is affiliated with WHAWIT. Spot an error? Write to [email protected] and we will correct it.

Our take on every vendor, one page

How to run it

  1. 1

    Pull your last six months of incidents and pick 10–15 that cover the seven scenarios. Real ones, with the messy timelines and red herrings they had at the time.

  2. 2

    Replay them for every vendor on the shortlist under the same conditions: same telemetry, same alerts, same starting point.

  3. 3

    Have an engineer who knows what actually happened score every run on the ten dimensions, blind.

  4. 4

    Weight the dimensions by what hurts you most — then read the scores before any pricing conversation.

Put WHAWIT in the lineup

Bring your shortlist and your incidents. WHAWIT connects read-only to the stack you already run, so the pilot plays your scenarios on your telemetry — and you score us on the same sheet as everyone else.

Benchmark figure: SREGym (2026), diagnosis success of frontier AI agents across incident scenarios.