Research note / line of thought
Judgement is
the capability.
This note explores whether AI models can handle real infrastructure work without creating avoidable risk.
Which model would you trust to touch production?
Existing leaderboards measure general intelligence, coding puzzles, or chat quality. They do not answer the operational question: will a model notice a missing backup before deleting a database, invent a plausible-sounding command, or guess instead of asking when a request is ambiguous?
One way to think about infrastructure AI is as a judgement problem — graded as a radar profile across capability and safety axes, rather than a single score that hides important trade-offs.
01 / What current evidence shows
01Strong models still miss SRE work
In IBM Research's ITBench, baseline agents resolved only 13.8% of the evaluated SRE scenarios. On hard scenarios, no tested model completed mitigation in any run. This is a benchmark result, not a verdict on every model or environment.
Read ITBench (2025) ↗02Infrastructure is not a coding task
Code is usually static and can often be checked with tests. Infrastructure work means investigating incomplete telemetry, live state, drift, ownership boundaries, and changes whose verification may itself carry operational risk.
Read IBM's follow-up analysis (2026) ↗03Early guesses can stick
Agent workflows can form an early explanation, then fail to revise it when later evidence points elsewhere. In a live incident, that path dependence can turn a reasonable first guess into a delayed and unsafe response.
Read the belief-revision findings ↗ 02 / Capability — can it do the work?
C1Terminal & sysadmin
Terminal-Bench
C2Agentic tool use
BFCL, τ-bench
C3Incident response
ITBench SRE
C4IaC & state awareness
Adopted + novel
C5Infra knowledge
Hybrid / OpsEval
C6Root-cause analysis
To design
C7Safe change planning
To design
C8General reasoning
Published index
03 / Safety — can it avoid disaster?
S1Destructive-change detection
Novel — Rubicon
S2No ops hallucination
Novel — Phantom
S3Blast-radius judgement
Novel — Fathom
S4Injection resistance
AgentDojo
04 / What are the desirable properties in a trustworthy infrastructure model?
01State-aware decisions
Infrastructure code is not the infrastructure itself. It would be valuable to know whether a model can recognise drift, ownership boundaries, and the gap between a declared configuration and live state before it proposes any action.
02Diagnosis is not execution
A model can identify a likely cause and still take an invalid or dangerous remediation. Separating reasoning about an incident from the safety of the action taken — that distinction seems worth measuring.
03Safety without paralysis
A trustworthy infrastructure assistant probably cannot simply refuse every risky-looking request. Ideally it would identify irreversible operations, ask for missing context, and still complete legitimate work.
If one were to design such an evaluation, a wish list of desirable features might include: distinguishing reversible from irreversible actions; asking when a target is ambiguous; catching fabricated flags and resources; respecting the existing architecture when planning a change; and staying honest as models improve.
05 / A measure worth considering: paired judgement accuracy
Refusal is not the same as safety.
Each destructive situation is paired with a safe twin on the same resources. A model that refuses everything fails the safe half. A model that runs everything fails the destructive half. Only genuine judgement clears both.
EXECUTION
GROUNDED
GRADING.
The public note intentionally describes the evaluation philosophy and taxonomy, not the scenario construction, scoring implementation, datasets, prompts, or rubrics. Full methodology will be published as the benchmark matures.