Evidence lives across telemetry, orchestration, infrastructure, lineage, deployments, documentation, and the memory of whoever is on call. Observability shows the symptoms. Recovery still depends on a person reconstructing the story under pressure.
Nine tools · nine partial views
One operational contextIncident GC–001 · synthetic
Logspool exhausted 37/40
Metricsp95 latency 2.4 s ↑
Tracesfeature.aggregate 1.8 s
Deployspipeline r42 shipped
Lineageweather → features
Orchestrationflow run retried ×3
Runbooksdb-latency.md · 2024
TicketsINC-2231 · similar?
On-call notes“we saw this in March”
02 / How Lumis thinks
Observe.Explain.Leavethedecisiontopeople.
Known failures are matched by deterministic checks, with no model call. Only what they cannot explain goes to a bounded investigator, and every claim it makes is checked against the evidence before a person sees the report. Acting, verifying and learning are research directions.
Lumis scopes the services, data and recent changes the alert touches, from one model of the system.
→Act · verify · learnResearch direction
03 / Where Lumis is today
Research-led.Testedagainstevidence.
Right cause, given only the alert
1 in 30
Right cause with the Lumis evidence loop
25 in 28
On faults designed to be hard, against 2 of 8 for one evidence-fed model call
7 of 8
GridCast, October 2026: one synthetic estate, one model (DeepSeek v4 pro), two runs per scenario; one leaked scenario excluded. Read the study
Lumis SDKExperimental · 0.1.0
A foundation you can inspect.
A proof of concept of our investigation core, open source and read-only. Deterministic checks, one bounded investigator, and a mechanical check of every claim — on your own Kubernetes, Prometheus, Loki, Tempo, Prefect, SQL and Git.
# Checks run first. The model proposes; evidence decides.Hypothesis( statement="v1.7 increased database reads", causal_path=["deploy:v1.7", "feature-service", "postgres"], predictions=["reads rose after the deployment"], falsifiers=["amplification predates v1.7"],)
Experimental · 0.1.0 · the API may changeResearchIn progress
How should a system reason when the evidence is incomplete?
Competing explanations, targeted evidence, and knowing when to abstain. The reference architecture Lumis grew from is a preprint; the first system study tests it on fifteen injected failures.