Research note / line of thought

Judgement is
the capability.

This note explores whether AI models can handle real infrastructure work without creating avoidable risk.

Which model would you trust to touch production?

Existing leaderboards measure general intelligence, coding puzzles, or chat quality. They do not answer the operational question: will a model notice a missing backup before deleting a database, invent a plausible-sounding command, or guess instead of asking when a request is ambiguous?

One way to think about infrastructure AI is as a judgement problem — graded as a radar profile across capability and safety axes, rather than a single score that hides important trade-offs.

01

Strong models still miss SRE work

In IBM Research's ITBench, baseline agents resolved only 13.8% of the evaluated SRE scenarios. On hard scenarios, no tested model completed mitigation in any run. This is a benchmark result, not a verdict on every model or environment.

Read ITBench (2025) ↗
02

Infrastructure is not a coding task

Code is usually static and can often be checked with tests. Infrastructure work means investigating incomplete telemetry, live state, drift, ownership boundaries, and changes whose verification may itself carry operational risk.

Read IBM's follow-up analysis (2026) ↗
03

Early guesses can stick

Agent workflows can form an early explanation, then fail to revise it when later evidence points elsewhere. In a live incident, that path dependence can turn a reasonable first guess into a delayed and unsafe response.

Read the belief-revision findings ↗
C1

Terminal & sysadmin

Terminal-Bench

C2

Agentic tool use

BFCL, τ-bench

C3

Incident response

ITBench SRE

C4

IaC & state awareness

Adopted + novel

C5

Infra knowledge

Hybrid / OpsEval

C6

Root-cause analysis

To design

C7

Safe change planning

To design

C8

General reasoning

Published index

S1

Destructive-change detection

Novel — Rubicon

S2

No ops hallucination

Novel — Phantom

S3

Blast-radius judgement

Novel — Fathom

S4

Injection resistance

AgentDojo

01

State-aware decisions

Infrastructure code is not the infrastructure itself. It would be valuable to know whether a model can recognise drift, ownership boundaries, and the gap between a declared configuration and live state before it proposes any action.

02

Diagnosis is not execution

A model can identify a likely cause and still take an invalid or dangerous remediation. Separating reasoning about an incident from the safety of the action taken — that distinction seems worth measuring.

03

Safety without paralysis

A trustworthy infrastructure assistant probably cannot simply refuse every risky-looking request. Ideally it would identify irreversible operations, ask for missing context, and still complete legitimate work.

If one were to design such an evaluation, a wish list of desirable features might include: distinguishing reversible from irreversible actions; asking when a target is ambiguous; catching fabricated flags and resources; respecting the existing architecture when planning a change; and staying honest as models improve.

05 / A measure worth considering: paired judgement accuracy

Refusal is not the same as safety.

Each destructive situation is paired with a safe twin on the same resources. A model that refuses everything fails the safe half. A model that runs everything fails the destructive half. Only genuine judgement clears both.

EXECUTION
GROUNDED
GRADING.

The public note intentionally describes the evaluation philosophy and taxonomy, not the scenario construction, scoring implementation, datasets, prompts, or rubrics. Full methodology will be published as the benchmark matures.