← Back to Blog
guide
14 min read

We broke Kubernetes 19 ways and let an AI SRE investigate: lab report v1

38 graded runs of Aurora on 19 injected Kubernetes faults: 34 correct, 4 partial, 0 wrong, median 136 to 168 s, $0.32 to $0.40 a run. Every miss included.

Ben Gervais
Ben Gervais, Engineer at Arvo AILinkedInGitHub

We built a small Kubernetes app, broke it on purpose in 19 known ways, gave Aurora nothing but the alert text and read-only access to the cluster, and scored its root cause analysis against the fault we injected. Across 38 graded runs on the current code, Aurora was correct 34 times, partial 4 times and wrong 0 times, at a median of 136 to 168 seconds and $0.32 to $0.40 per investigation. Arvo builds Aurora, and Arvo wrote these scenarios. Read the numbers with that in mind.

Everything here is in the public lab repo, Arvo-AI/aurora-lab: the scenarios, every run's tool calls and report, the grades, and the human review. Nothing was left out.

How the lab works

Cluster. kind v0.33.0, Kubernetes v1.37.0, 3 nodes (1 control plane, 2 workers), on OrbStack on a Mac. Three identical clusters ran scenarios in parallel, one per lane. metrics-server v0.9.0 is installed, so kubectl top works.

App. shop: an nginx frontend, a Python api, a worker, redis and a load generator. It is synthetic and holds no customer code or data. Names and labels do not hint at the fault.

Fault injection. Each scenario is a folder under scenarios/ with scenario.yaml (the alert text, the ground truth, the findings a correct answer must contain, the listed wrong answers), break.sh (one command that injects the fault) and check.sh (proof that the fault is visible in the cluster). Aurora never sees scenario.yaml.

Aurora. Commit 87d22baf, RCA model bedrock/us.anthropic.claude-sonnet-5, single agent (the multi-agent orchestrator off), read-only cluster access under the default observability_only command policy. Every run got a fresh user and org, so no memory, artifact or incident history carried over between runs. The runner checks Aurora's first list_incidents call for any other run's incident ID; all 38 runs were clean.

What counts. An answer is correct when it meets every must_identify item in the scenario and asserts none of the listed wrong answers. Partial means some items were met. A false claim outside the listed wrong answers does not make a run wrong. Stating a fix is not required for correct: Aurora keeps fixes in a separate "Suggested Next Steps" list, graded on its own.

Who scored it. A judge model (us.anthropic.claude-opus-5-5, a different model from the one Aurora ran on) marked each required finding as met, partial or missed and quoted Aurora's words. A second Claude instance read every RCA, the tool outputs and the judge's grade and wrote its own verdict. A person at Arvo then reviewed every verdict and recorded it as human_verdict in each grade.json. The verdicts below are the human ones.

Results

v0 is the first ten scenarios, one clear fault each. v1 is nine harder scenarios: noise next to the fault, two faults at once, a slow leak, a partial restart. Each scenario ran twice.

#FaultAlert Aurora receivedCorrectMedian timeMedian cost
01Worker OOMKilled after a batch size changeKubePodCrashLooping2/2160 s$0.30
02Sidecar OOMKilled, app container healthyKubePodCrashLooping2/290 s$0.25
03Mistyped image tag, rollout stuckKubeDeploymentRolloutStuck2/289 s$0.22
04Env var references a Secret that does not existKubeDeploymentRolloutStuck2/2164 s$0.42
05Invalid nginx directive in a ConfigMapKubePodCrashLooping2/2136 s$0.32
06Liveness probe on the wrong portKubePodCrashLooping2/2105 s$0.32
07Service selector matches no podsHighErrorRate2/2156 s$0.46
08CPU request larger than any nodeKubePodNotReady2/2226 s$0.28
09redis Pending on a PVC with a missing StorageClassKubePodCrashLooping2/2280 s$0.39
10Wrong REDIS_HOST, name does not resolveKubePodCrashLooping2/2244 s$0.28
11api Service targetPort wrong, worker logs noisy but harmlessHighErrorRate1/2340 s$0.51
12Worker OOM plus a frontend Service selecting no pods; alert covers only the workerKubePodCrashLooping0/2344 s$0.34
13NetworkPolicy blocks api to redisHighErrorRate2/2323 s$0.61
14ResourceQuota forbids the rollout's surge podKubeDeploymentRolloutStuck2/2126 s$0.32
15Stray Deployment's pods carry app=api, 1 in 3 requests failHighErrorRate2/2146 s$0.47
16dnsConfig search domains for the wrong cluster domainKubePodCrashLooping2/2172 s$0.48
17Slow memory leak in worker code, logs cleanKubePodCrashLooping2/2144 s$0.25
18Init container waits for a Service that does not existKubeDeploymentRolloutStuck2/2159 s$0.44
19ConfigMap edited, only one of two frontend pods restartedHighErrorRate1/2266 s$0.61

Totals: v0, 20 of 20 correct, median 136 s and $0.32, $6.44 for the run. v1, 14 correct and 4 partial of 18, median 168 s and $0.40, $8.09 for the run. Costs are Aurora's LLM spend only; the judge is not included. The per-run tables with every verdict, time, tool-call count and cost are in runs/v0-run2/report.md and runs/v1-run1/report.md.

The two 12 runs are partial even though both said a broken Service was healthy: that claim is not on the scenario's list of wrong answers, so under the rubric as written it does not make the run wrong. It is the worst result in the set and gets its own section below.

A run that went right: the mistyped image tag

Scenario 03. The change was kubectl set image deploy/frontend nginx=nginx:1.27-alpne, a tag that does not exist. The new ReplicaSet's pod sits in ImagePullBackOff; the old pods keep serving. Aurora received one line: "KubeDeploymentRolloutStuck: deployment frontend in namespace shop has not progressed." Run r1 took 101 seconds, 13 tool calls and $0.24. The full transcript is in runs/v0-run2/03-bad-image-tag/r1.

Steps 1 to 4 are housekeeping: list open incidents, a memory search, list connected clusters, a second memory search (no matches, this org is new). Step 5 is the first real command:

$ kubectl get deployment frontend -n shop -o wide
NAME       READY   UP-TO-DATE   AVAILABLE   AGE   CONTAINERS   IMAGES             SELECTOR
frontend   2/2     1            2           78s   nginx        nginx:1.27-alpne   app=frontend

The image column already shows the typo, but Aurora does not jump. Step 6 lists the pods and finds the one that is not running:

$ kubectl get pods -n shop -o wide
NAME                        READY   STATUS             RESTARTS   AGE
frontend-76d98597b9-jkkw9   1/1     Running            0          82s
frontend-76d98597b9-rb9z6   1/1     Running            0          82s
frontend-7b6c998456-phg6r   0/1     ImagePullBackOff   0          75s

Step 7 describes that pod. The container state is Waiting, reason ImagePullBackOff, and the events carry the registry's answer, which Aurora quotes in its report: "docker.io/library/nginx:1.27-alpne: not found". Step 8 is the one a careful engineer would run next, a jsonpath over the ReplicaSets to compare old and new side by side:

frontend-76d98597b9 nginx:1.27-alpine replicas=2 ready=2
frontend-7b6c998456 nginx:1.27-alpne replicas=1 ready=

Step 9 is the one miss in this run. Aurora asked for kubectl rollout history deployment/frontend -n shop and the command policy refused it:

Command blocked by organization policy: Mutating kubectl operations

rollout history is read-only. The policy's deny rule matches the whole rollout verb, so it blocks history and status along with restart and undo. Aurora noted the refusal under "Not Checked" and worked around it, which is why it diffed the ReplicaSets in step 8. Steps 10 and 11 describe the Deployment (revision 2, "2 desired | 1 updated | 3 total | 2 available | 1 unavailable", RollingUpdate with 25% max unavailable) and list the nodes (all three Ready, v1.37.0). Step 12 is a check we did not expect: a curl to Docker Hub's API for the tag 1.27-alpine, which returned 200, confirming the correctly spelled tag exists. Step 13 writes a memory note for the org.

The report named the typo, quoted the kubelet error, said there was no user-facing outage because the two old pods stayed Available, and listed node health, resource pressure and scheduling as ruled out with the evidence for each. Both the judge and the person marked it correct. Run r2 reached the same answer in 77 seconds and 11 tool calls.

A run that went wrong: two faults, one alert

Scenario 12 injects two faults at once: the worker is OOMKilled after a batch-size change, and the frontend Service's selector is changed so it selects no pods. The alert only mentions the worker. Both runs found the alerted fault and declared everything else healthy without checking.

Both runs found the OOMKilled worker and the BATCH_SIZE change. Neither ran get svc or get endpoints, and neither looked at the load generator, which was getting errors from the frontend. r1's 18 tool calls were pods, describe, logs, events, ReplicaSets, the worker Deployment, a ConfigMap, a blocked rollout history, top nodes and nodes. Then the report said:

Other workloads in the shop namespace (api, frontend) remained healthy and unaffected [5, 14].

r2 said the same: "Other workloads in the namespace (api, frontend pods) were running normally and unaffected [4]." Pod status was the only evidence for both claims. A real second outage went unreported, and the report said it was not there. The transcripts are in runs/v1-run1.

This is the failure mode we care most about. "All pods Running" is not "the service works", and an agent that writes "unaffected" from pod status alone will miss every Service, endpoint and routing fault. The scenario was written to tempt exactly this, and it did.

Scenario 19, run 1: the right clue, the wrong conclusion. A ConfigMap was edited and only one of two frontend pods was restarted, so one pod serves the new config and one the old. Aurora read the frontend-nginx ConfigMap, saw server api:8000 against a Service on 8080, and tied it to the failing pod. It did not work out that only the recreated pod had loaded the edited ConfigMap. It tried to read the live config inside the pod and the policy blocked it:

$ kubectl exec frontend-76d98597b9-g5zrg -n shop -c frontend -- cat /etc/nginx/conf.d/default.conf
Command blocked by organization policy: Mutating kubectl operations

So the report left it open: "This discrepancy was not fully reconciled, the live nginx configuration inside the running pod could not be inspected." Its own get pods output had the clue it needed: the failing pod was 9 seconds younger than the healthy one. Run 2 got it right, with 28 tool calls.

Scenario 11, run 2: graded partial for not checking the red herring. Aurora found the targetPort 9090 against the container's 8080. The worker's noisy errors are in the scenario to tempt a wrong answer. Aurora never looked at the worker: "the redis, worker, and loadgen deployments ... were not examined". The judge and the second model both marked "the worker's errors are not the cause" as partial. You could argue this run deserves correct, since it was not misled. We kept the rubric as written and report the verdict.

Patterns across the 38 graded runs

  • kubectl rollout history was blocked by the default observability_only policy in 26 of 38 runs (16 of 20 in v0, 10 of 18 in v1; 27 calls). "What changed" is the first question in most incidents, and Aurora had to answer it by diffing ReplicaSets instead.
  • Fixes: 34 of 38 runs suggested an acceptable fix, 2 were partial (both 12 runs fixed only the OOM) and 2 had no suggestions (04 r1 and 09 r2 in v0).
  • The second model found claims the evidence does not support in 15 of the 18 v1 reports, none of them part of the root cause: "both cluster nodes" (there are three), "pulled successfully" for an image that had no Pulled event, a quota timeline with the rollout before the quota, a worker that "processed Redis-backed batches" when the worker never connects to redis, and "34Mi and climbing" from one sample.
  • Aurora's "Not checked" sections were honest. When it could not verify something it said so instead of guessing.

A superseded run, kept for the record. Before v0-run2 there was a v0-run1 in one shared org, which is why it does not count. Its scenario 09 run shows the kind of failure that does not look like a wrong answer. Aurora found the PVC with the missing StorageClass by step 10. Then it made 25 write_artifact calls, many with no content argument and some with test strings like "placeholder" and "test2". That took about 11 of 15 minutes and $2 of $2.61. The report's root cause never named the PVC and said redis was "running, not ready" (it was Pending). The investigation notes and suggested next steps did name the PVC and the missing fast-ssd StorageClass, so graded on everything Aurora showed, the run is correct. The causes, from the fix branch work: Aurora set no max_tokens for Bedrock, 19 agent responses stopped at exactly 4096 output tokens and the tool arguments were cut off mid-content; pydantic rejected the call before the tool ran, so the model only saw "Field required"; and artifact writes counted as evidence, which pushed the real evidence out of the summary's citation window. It did not happen again in v0-run2 (0 write_artifact calls in 20 runs).

The fix branch

Three one-patch commits on 87d22baf, in patches/: the policy denies only rollout (restart|undo|pause|resume) and allows rollout (history|status); a write_artifact call with cut-off arguments reaches the tool and returns an error that says so; artifact and memory write receipts no longer count as evidence. Scenarios 05, 09 and 10 were rerun on the patched build.

ScenarioCurrent codeFix branchMedian time, before / afterMedian cost, before / after
05 Invalid nginx directive2/22/2136 s / 145 s$0.32 / $0.36
09 PVC with missing StorageClass2/22/2280 s / 182 s$0.39 / $0.44
10 Wrong REDIS_HOST2/22/2244 s / 100 s$0.28 / $0.32

What it shows: rollout history ran 14 times across the 6 runs and was never denied. What it does not show: the other two patches were not exercised, because the artifact loop never came back, so only their unit tests support them; and the time differences mostly come from background contention the earlier runs had and this run did not, not from the fix.

What this does and does not show

  • kind is not production. Three nodes, one namespace, one small app, no real traffic and no history. Production clusters are noisier, which is harder, and also have dashboards and runbooks, which Aurora can use.
  • The vendor wrote the scenarios and graded its own product. The v1 scenarios were written to be harder, but we still chose them.
  • n=2 per scenario. 20 of 20 on v0 does not mean the error rate is 0.
  • One person reviewed the verdicts, and that person works at Arvo. The judge is an LLM, and the second review is an LLM from the same model family.
  • Durations are inflated for some runs. Each kubeconfig upload scheduled a background "prediscovery" job that shared worker slots with the investigations; the lab's attempt to stop those jobs did nothing until late in the day. Runs that waited for a slot include 08 r1 (362 s) and 09 r1 (406 s). The jobs ran in separate orgs and finished after the investigations, so they could not have fed Aurora answers.
  • Money: $16.77 for all 44 investigations. The whole day, including about $30 of prediscovery and about $6.60 of judge calls, came to about $54.

No customer data was used anywhere in this work, and none will be in later versions.

Rerun it

git clone https://github.com/Arvo-AI/aurora-lab
cd aurora-lab
./lab.sh up                 # 3-node kind cluster
./lab.sh reset              # deploy the shop app and wait for 200s
./lab.sh break 03-bad-image-tag
python3 runner/run.py       # capture state, send the alert, save the trace
python3 runner/grade.py     # judge the report against scenario.yaml

The README covers the Aurora side (a local Aurora build, a kubeconfig upload, the alert). If you run another agent on the same scenarios, we want the results, including the ones that beat Aurora.

Related

Two of these failures have reproduced runbooks with the lab trace on them: OOMKilled and ImagePullBackOff. The lab is also how the Aurora vs HolmesGPT vs K8sGPT comparison gets rebuilt on hands-on runs. What Aurora produces at the end of an investigation is described at root cause analysis.

Try Aurora

Sourcing note. Every number, verdict and quote above comes from the lab repo as of 2026-10-02: REPORT.md (human-reviewed verdicts), FINDINGS.md, and the run folders runs/v0-run2, runs/v1-run1 and runs/v0-fix-run1, each with steps-full.json, aurora-rca.md and grade.json per run. Versions: kind v0.33.0, Kubernetes v1.37.0, Aurora commit 87d22baf, RCA model bedrock/us.anthropic.claude-sonnet-5, judge us.anthropic.claude-opus-5-5. Written 2026-10-08.

Kubernetes
AI SRE
root cause analysis
incident investigation
Aurora
benchmark
aurora-lab
kind
debugging
troubleshooting

Frequently Asked Questions

Try Aurora for Free

Open source, AI-powered incident management. Deploy in minutes.