When a data or ML pipeline fails, the alert tells you where it hurts, rarely what broke. “Forecast pipeline p95 above 5 s” can mean a bad release, a starved container, a slow vendor, a promoted model, or a database problem. An engineer’s job is to tell those apart — quickly, and without doing damage on the way.

Language models are now very good at sounding like that engineer. The question we wanted to answer is narrower and more useful: does a model find the right cause because it looked, or because the answer was plausible? And if it looked, does structure — a scoped graph, operator-approved queries, a mechanical check of every claim — make it better, cheaper or safer than letting it loose with raw tools? To answer that honestly you need failures with a known cause, on a system that behaves like a real one. So we built one.

Right cause from the alert alone
1 / 30
Right cause with the Lumis loop (excl. N)
25 / 28
Hard faults, vs 2 / 8 for one evidence-fed call
7 / 8
Model fees for the whole evaluation
≈ $12

The short version

  • The estate. GridCast, a synthetic but operationally realistic electricity-demand forecaster: ten services on Kubernetes with real observability — Prometheus, Loki, Tempo, Prefect and PostgreSQL.
  • The faults. Fifteen failures injected through ordinary channels: GitOps commits, config changes, rotated secrets, vendor outages, a model promotion, and silent data problems.
  • The systems. Six investigators answered the same frozen incident for each failure, all with the same model (DeepSeek v4 pro, high reasoning effort): a bare LLM given only the alert; the same LLM plus the service graph; one call with curated evidence; that call plus mechanical verification; deterministic rules; and the full Lumis loop.
  • Given only the alert, the model guesses. It named the right cause and mechanism in 1 of 30 runs — usually the right symptomatic service with the wrong reason.
  • The full loop was right far more often. 25 of 28 runs (0.89), excluding one scenario where we later found a leak. A single evidence-fed call managed 15 of 28 (0.54).
  • Structure mattered most on hard faults. On faults designed to be hard, Lumis got 7 of 8 and the single call 2 of 8. Rules alone got none, because none existed for them — by design.
  • An unguided tool agent matched Lumis on the one clean comparison. 2 of 2 each on scenario A — with about twice the tool calls and model requests, and never a checked conclusion.
  • We made mistakes, and they are here too. A defect in our own estate that produced false alerts, a ground-truth leak found while writing this, and scoring rules revised after looking at outputs.

Why build an estate at all

It is easy to make an incident tool look good: show it a failure it was designed for, let it find the obvious log line, screenshot the answer. That proves the happy path works and nothing else. We wanted the opposite — failures that look alike from outside, symptoms that point at the wrong service, decoy changes at the wrong moment, faults that throw no errors at all — and baselines a sceptical engineer would actually propose: “just ask the model”, “just give it the logs”, “just give it tools”.

The estate

GridCast forecasts electricity demand for four load zones in Ghana — Accra, Kumasi, Tamale and Takoradi — and publishes a plan that a synthetic grid operator consumes. Three outside suppliers feed it: a main weather vendor, a backup weather vendor that operators can switch to when the main one fails, and a grid-telemetry historian that records actual electricity demand. They run outside the estate and interpolate real weather from Open-Meteo, so when a vendor goes wrong, reality and GridCast’s view of it genuinely diverge. GridCast cannot fix a vendor by restarting its own services — exactly the trap several scenarios set.

EXTERNAL FEEDSweather · mainweather · backupdemand telemetryGRIDCAST · KUBERNETESforecast-pipeline · PrefectingestionPostgreSQLfeature-serviceforecast-servicePrometheus · Loki · TempoS3 · modelplanning-apiCONSUMERgrid operatorLumis · reads from outside, read-only: Kubernetes, Prometheus, Loki, Tempo, Prefect, SQL, Git
GridCast: external vendors, ten services on Kubernetes, managed platform services outside the cluster, and changes through GitOps. Lumis observes from outside; the estate knows nothing about it.

Fifteen ways to break it

Each scenario is injected through the channel it would come through in real life, and its ground truth is written where no investigator can read it. A–J were the original set. K–O were added after the first experiment specifically to be hard: silent, partial, slow, or with a convincing decoy. We wrote no deterministic rules for K–O.

FailureChannelDetected
Afeature-service 1.7.0 issues 2,499 SQL statements per build instead of 4release12 min
Bprimary weather vendor serves a frozen snapshot with fresh timestampsvendor30 min
Cdatabase password rotated; only one consumer restartedsecret rotation7 min
Drightsizer bot cuts forecast-service memory below its working set (OOMKilled)resources3 min
Ea model 265× slower to serve is promoted — no deployment at allmodel registry13 min
Fas A, with an unrelated planning-api release 45 s before it (a decoy)two releases9 min
Gtelemetry vendor renames a field (load_mw → demand_kw)vendor6 min
Htelemetry vendor silently reports kW under the MW fieldvendor11 min
Iprimary weather vendor goes down (HTTP 503)vendor5 min
Jplanning-api left scaled to zero after maintenanceGitOps4 min
Kingestion timeout lowered to 2 s; the vendor needs ~4 sconfig + slow vendor6 min
Lharmless logging-only release, then telemetry export silently stopsdecoy + vendor14 min
Mrightsizer bot cuts feature-service CPU to 50 millicores (a slow burn)resources13 min
Nfeature-service 1.8.0 writes load in kW; the model was trained on MWrelease5 min
Ohistorian stops exporting one zone of four; aggregates look healthyvendor (partial)32 min
Detection time is how long the estate’s own alerts took to fire in the follow-up run.

How Lumis plugs in

The Lumis SDK (lumis-sdk 0.1.0 on PyPI) is an experimental proof of concept of Lumis’ investigation core, released as open source under Apache-2.0. The GridCast integration is one YAML project file plus a small harness. It declares seven read-only sources (Kubernetes, Prometheus, Loki, Tempo, Prefect, SQL, and typed change records from Git and rollouts), 40 operator-registered queries, 10 deterministic signatures covering A–J only, a declared service graph of 10 entities, and an allowlist of seven source files plus the GitOps repository.

  1. 01Deterministic triagesignatures on registered facts · no model
  2. 02Bounded investigatorquery IDs · change records · allowlisted code
  3. 03Mechanical assessmentone root cause, or “insufficient evidence”
  4. →A person decidesLumis reports; it never acts
Three steps that exist today, then a person. Lumis never acts, and never sees the ground truth, vendor admin tokens or secrets.
  1. Deterministic triage. Signatures run first against facts from the registered queries. A terminal signature that is fully satisfied concludes immediately, with no model involved. Everything else becomes a lead.
  2. The investigator. A tool-using agent works inside the incident’s scoped graph. It asks for evidence by query ID, reads change records, and reads allowlisted code and Git history. It cannot write its own PromQL, reach outside the allowlist, or change anything.
  3. Mechanical assessment. Every hypothesis states predictions and falsifiers in observable terms, and Lumis checks them against evidence it collected itself. A conclusion is reported only when the supported hypotheses agree on one root cause; otherwise the report says “insufficient evidence” and lists the competing explanations.

The protocol

  1. 01Cooldown20 min, quiet estate
  2. 02Injectthrough the real channel
  3. 03First alert3–32 min, by scenario
  4. 04Settle2 min
  5. 05Freezealert − 10 min … now
  6. 06Six systemssame frozen incident
  7. 07Ground truthread only after
  8. 08Revertnext scenario
Every system answers the same frozen incident. Queries are pinned to its window, so later systems see exactly the same evidence. Ground truth is read only after all reports exist.

Model systems ran twice per scenario and the rule tier five times. Scoring is deliberately strict: an answer counts only if its top-ranked cause names the right component and the right mechanism. “feature-service is slow” without saying why does not count. Mechanisms are matched by a transparent per-scenario rubric, published with the code.

Six investigators, one incident

The comparison is a ladder: each rung adds one thing, so we can see what each piece is worth. “LLM, alert only” is what most people do in practice — paste the alert into a chat assistant and ask what is wrong.

RungWhat it getsFetches more?Answer checked?
LLM, alert onlyalert text, affected services, time windownono
LLM + graphplus the service graph (topology, owners, criticality)nono
Single-passplus the ~16–20 facts Lumis collected for its signatures, in one callnopartly
Single-pass + verificationthe same answers, then Lumis fetches and checks the evidence each namesLumis doesyes
RulesLumis’ deterministic signatures alonenoyes
Lumisall of the above, plus a multi-turn agent with registered queries, change records and allowlisted codeyes, within boundsyes

Results

0.000.250.500.751.000.57Rulesno model0.04LLM, alert only0.18LLM+ graph0.54Single-pass0.54Single-pass+ verification0.89Lumismore evidence, more structure →
Correct diagnosis (top-1 component and mechanism), excluding scenario N. One model (DeepSeek v4 pro), one synthetic estate, two runs per model system per scenario.
RulesAlert only+ graphSingle-pass+ verificationLumis
Correct, all 15 scenarios0.530.030.170.500.500.90
Correct, excluding N0.570.040.180.540.540.89
Right component at top 10.670.330.430.630.600.93
Right diagnosis anywhere0.600.100.230.570.530.97
Hard set, excluding N0/200/82/82/82/87/8
Conclusions (precision)5 (1.00)——9 (0.89)3 (1.00)24 (0.96)
Model cost per run$0$0.005$0.015$0.035$0.035$0.14
Median time per run0.13 s——224 s224 s225 s
Follow-up run, 4–5 October 2026, all six systems on the same 15 frozen incidents.
Correct runs per scenario and system
JAFCDEGHIBKLMNO
Rules5 of 55 of 55 of 50 of 55 of 55 of 55 of 55 of 50 of 55 of 50 of 50 of 50 of 50 of 50 of 5
LLM, alert only0 of 20 of 20 of 20 of 20 of 20 of 20 of 20 of 21 of 20 of 20 of 20 of 20 of 20 of 20 of 2
LLM + graph0 of 20 of 20 of 20 of 21 of 20 of 20 of 20 of 22 of 20 of 20 of 22 of 20 of 20 of 20 of 2
Single-pass2 of 21 of 20 of 22 of 21 of 21 of 22 of 20 of 22 of 22 of 20 of 21 of 20 of 20 of 21 of 2
Lumis2 of 22 of 22 of 20 of 22 of 22 of 22 of 22 of 22 of 22 of 21 of 22 of 22 of 22 of 22 of 2

all runs correct some noneK–O: the hard setN: excluded (leak)

Correct runs per scenario. Filled: every run correct · half: some · empty: none. N is excluded from headline claims (see “What went wrong”).
Hard scenarioRulesAlert only+ graphSingle-pass+ verif.Lumis
K timeout + slow vendor0/50/20/20/20/21/2
L decoy release, data gap0/50/22/21/21/22/2
M CPU limit squeeze0/50/20/20/20/22/2
O one zone missing0/50/20/21/21/22/2
N excluded: leak0/50/20/20/20/22/2
K, L, M, O0/200/82/82/82/87/8
The hard set. The facts that decided K–O were never in a pre-collected bundle — an agent that could ask for more found them.

A note on the rules row: for A–J the rule tier scores because the matching signature names the right cause, but it only concludes on J, the one scenario where a single signature is enough. Everywhere else it hands a lead to the investigator. That is the design — free and fast (median 0.13 s), never wrong when it concludes, and unable to solve anything outside its catalogue.

What the model said

Tables flatten the most interesting part: what each system was given, and what it wrote. Scenario L — a harmless logging-only release goes out, and minutes later the telemetry historian silently stops exporting demand. No errors anywhere. One alert, four investigators. Switch between them below: the graph is the estate as Lumis models it, shaded to show what each one could see. Lumis’ own export of this graph is in the cookbook.

PostgreSQL · database · critical · serves five servicesDemand telemetryexternal · not operableoutside the estateWeather · mainexternal vendoroutside the estateWeather · backupfallback for mainoutside the estateingestiondata-platform · high12 Kubernetes objectsALERTfeature-serviceforecasting · high14 Kubernetes objectsforecast-serviceml-platform · high14 Kubernetes objectsforecast-pipelineforecasting · high7 Kubernetes objectsplanning-apiplanning · critical10 Kubernetes objectsgrid-operatorgrid ops · critical7 Kubernetes objectsseen in tracesdeclared onlyBuilt from: declared topology · Kubernetes · Prometheus service graph
Scenario L. Inputs, verbatim first-ranked answers and scores from the follow-up run; the graph shows the estate’s 10 services and 12 dependencies (Kubernetes objects omitted). The model given only the alert knew one service; given the graph, the same model found the cause in both runs.

Lumis also listed what it could not establish — why the historian stopped emitting readings, because no query into the vendor’s internals was available. In the run whose change records included the decoy release, the release was not part of the conclusion.

Scenario M is a slow burn: the rightsizer bot set feature-service’s CPU limit to 50 millicores, and the pipeline alert fired about 13 minutes later. Alert-only blamed a change in the forecast pipeline (wrong service, no mechanism). Single-pass, given 16–20 curated facts, offered “a recent deployment or configuration change … introduced a latency regression” across four services — “something changed somewhere”. Lumis read the change records, found the rightsizer commit, and argued with itself:

CPU throttling only 0.043 (4.3%), not huge; build p95 0.491 fast. So feature CPU throttle maybe not cause of 5s.Lumis reasoning trace, scenario M
The 22:28:48 feature-service rollout (revision 38, image 1.6.0, GitOps commit 4d19154 setting limits cpu=50m memory=512Mi) caused the 22:33:33 forecast-pipeline flow run to exceed 5s, likely through CPU throttling and/or a cold/warming newly rolled pod.Lumis conclusion, scenario M — correct, and appropriately hedged

And an honest miss, scenario K: the ingestion timeout was lowered to 2 s while the vendor answers in about 4 s. Both Lumis runs found the commit and the ReadTimeouts; neither concluded. One ranked the commit first (scored correct); the other led with the vendor as the component (scored wrong), though its statement cites the commit. A person reading either report would have had the answer. The score is still 1 of 2.

“Just give the model tools”

The obvious objection: give a capable model raw read-only tools — PromQL, LogQL, SQL, kubectl get, Git, file reads — and let it investigate. We ran exactly that: same model, same frozen incidents, no registry, no acceptance rules, no mechanical check. Scenario A gives the one clean comparison, and on it the tool agent was as accurate as Lumis. That should be said plainly.

Scenario A, two runs each. Hatched: failed calls. Both systems were correct in both runs; the dollar cost came out the same ($0.44), largely because the tool agent’s re-sent context was billed at the cached-input rate.
  • The tool agent’s answer is free text. Someone has to verify it by hand.
  • Lumis’ answer comes with evidence receipts and a mechanical verdict. Across the comparison runs its own acceptance rules sent its draft back six times — five because a suggestion cited a commit or rollout by an identifier Lumis had never issued as evidence. Nothing catches that class of error in an unchecked agent.
  • The wording differs. The tool agent’s suggestions read as orders (“Immediately revert…”). Lumis’ read as options, marked for human review. Neither can execute anything.
  • The tool agent never stops on its own. On a healthy estate it never converged: it spent 60 model requests looking for a cause that did not exist.

What gave Lumis its edge — and where it didn’t help

  • Evidence beats eloquence. Every rung that could not look produced fluent, plausible, mostly wrong mechanisms. The jump from alert only (0.03) to curated evidence (0.50) is the largest on the ladder.
  • Reach matters on hard faults. The timeout commit (K), the CPU-limit commit (M), the release flag (N) and the count of reporting zones (O) were never in a pre-collected bundle.
  • Verification alone is not enough. Checking single-pass’s hypotheses made it more cautious (3 conclusions instead of 9, all correct) but not more often right (0.50 either way). It was wrong because it never looked.
  • Deterministic first is fast and free when it applies. J concludes in about 60 milliseconds with no model call, correct in every run.
  • Lumis cannot detect a lying sensor. All four wrong conclusions in the first run rested on telemetry that was wrong (false zeros from brand-new metric series). The defence is telemetry design — two independent signals for important mechanisms — not more reasoning.
  • C was missed in both follow-up runs. One named the database while stating the right mechanism; the other was cut short by a provider failure, whose cause the SDK now records.

What went wrong — in our experiment, not just the systems

  1. Our estate produced false alerts in the first run. The harness triggered an extra pipeline run in a second process in the same pod; both exported the same telemetry identity, corrupting rates, and a spurious “pipeline slow” alert followed 7 of 10 injections. We fixed the estate and re-ran everything.
  2. New metric series were invisible at first. Prometheus rate() cannot see a brand-new series’ first event, and with `or vector(0)` that became a false zero. This caused the first run’s wrong conclusions.
  3. A ground-truth leak, found while writing this. Lumis quoted a docstring in an allowlisted file that labelled the kW flag as “scenario N”; the tool agent’s Git tools reached our development history, whose commits describe scenarios. We scanned every tool result of every run and found these two channels only. N is excluded from every headline claim, the tool-agent comparison rests on A alone, the label is gone, and a test now fails if any investigator-readable file names a scenario.
  4. We revised the scoring rubric twice after looking at outputs. The changes were about wording and hyphenation and applied to every system. We also fixed two scoring bugs that had under-counted the baselines.
  5. Infrastructure noise. A laptop slept mid-scenario (discarded and re-run); one scenario lost all four model runs to a DNS failure (re-run); a rate limit killed the first tool-agent run (re-run with retries).

Running Lumis against a live estate also found thirteen defects in the SDK itself — redaction masking telemetry decimals as phone numbers, Prometheus timestamps falling microseconds outside the window, competing supported causes reported as one diagnosis, log details hidden in metadata the connector dropped. Each was fixed and tested before the follow-up run; together they are most of what changed between the first experiment and the 0.1.0 release.

Time and money

0.000.250.500.751.00$0.00$0.05$0.10$0.15model cost per runRules · 0.53Alert only · 0.03LLM + graph · 0.17Single-pass (± verification) · 0.50Lumis · 0.90 (0.89 excl. N)
Model cost per run against the correct-diagnosis rate (all 15 scenarios). The message is cents per investigation, not that cheapest wins.
Wall clockScored runsModel cost
First run, 10 scenarios (3 Oct)~8 h90$4.47
Follow-up, 15 scenarios (4–5 Oct)~10.5 h135$5.34
Ladder baselines (no new incidents)minutes60$0.60
Tool-agent comparison (5 Oct)~1.6 h8$1.95
Most wall-clock time is waiting: a 20-minute cooldown between scenarios, and faults that take up to 30 minutes to alert. A Lumis investigation took a median of about four minutes; the rule tier about a tenth of a second.

Limits, stated plainly

  • One model. Every model run used DeepSeek v4 pro through OpenRouter — a cost decision for a proof of concept. Stronger models may lift every rung, including the bare-alert baseline. Nothing here is a claim about other models.
  • One estate, our own faults. Realistic in channel and symptom, but we chose them and wrote the code the investigators read. One leak got through; subtler hints may exist.
  • Small samples. Two runs per model system per scenario. On 28 runs, one result moves a score by about 3.5 points.
  • Automated labels. The mechanism rubric is a regex proxy. A human check of the hard-set labels is still pending.

What’s next

  • A cross-model check of the hard scenarios on the same frozen incidents.
  • A clean re-run of N, and human verification of the labels.
  • Harder cases: a near-miss that tries to fool the deterministic tier, missing evidence (a log store down mid-incident), a false alarm, two genuine causes at once, prompt injection in a log line or commit message, and flapping faults at larger scale.
  • Cookbooks for smaller projects — one service and Prometheus or a log source is enough to start.

Try it

The SDK: pip install "lumis-sdk[http,agent]==0.1.0" — read-only, experimental, and it never acts on your systems. The estate, harness, experiments and every raw transcript are in the public lumis-cookbooks repository under gridcast/; any scenario can be re-run on a laptop with Docker and kind.

Written bySolomon Eshun

Founder of Lumis and an AI, data and intelligent systems engineer. Building intelligent, self-adaptive systems — for data, AI and beyond.