Intelligence for systems that cannot afford to guess.

Lumis — the film

A one-minute film with music and sound design, no narration. On-screen text carries the story; captions are available. Illustrative — synthetic incidents, not a live product session.

Read the film transcript

Everything runs on systems: pipelines, models, infrastructure. When they fail, the evidence is everywhere, and engineers rebuild the story by hand, under pressure, every time. Lumis builds one model of the whole system. Known failures are matched deterministically, in seconds, with no model call. Only when checks fall short does one bounded investigator inspect, probe, and test competing explanations — evidence, not guesses. A person decides, and recovery is verified, never assumed. Every incident makes the next one easier, and recurring fixes become deterministic rules once a person approves. Today: Data & AI systems. Next: cloud, edge, industry. The goal: intelligent systems that adapt, learn, and heal themselves. Intelligence for systems that cannot afford to guess. Lumis — starting with Data & AI. Follow the research.

Scroll to explore

01 / The problem

Complex systems fail in ways no single tool can explain.

Evidence lives across telemetry, orchestration, infrastructure, lineage, deployments, documentation, and the memory of whoever is on call. Observability shows the symptoms. Recovery still depends on a person reconstructing the story under pressure.

Nine tools · nine partial views
One operational contextIncident GC–001 · synthetic
Logspool exhausted 37/40
Metricsp95 latency 2.4 s ↑
Tracesfeature.aggregate 1.8 s
Deployspipeline r42 shipped
Lineageweather → features
Orchestrationflow run retried ×3
Runbooksdb-latency.md · 2024
TicketsINC-2231 · similar?
On-call notes“we saw this in March”

02 / How Lumis thinks

Observe. Explain. Leave the decision to people.

Known failures are matched by deterministic checks, with no model call. Only what they cannot explain goes to a bounded investigator, and every claim it makes is checked against the evidence before a person sees the report. Acting, verifying and learning are research directions.

See how it did on fifteen failures
Stage 01 of 05: Alert
  1. A symptom arrives.

    Lumis scopes the services, data and recent changes the alert touches, from one model of the system.

  2. →Act · verify · learnResearch direction

03 / Where Lumis is today

Research-led. Tested against evidence.

Right cause, given only the alert
1 in 30
Right cause with the Lumis evidence loop
25 in 28
On faults designed to be hard, against 2 of 8 for one evidence-fed model call
7 of 8

GridCast, October 2026: one synthetic estate, one model (DeepSeek v4 pro), two runs per scenario; one leaked scenario excluded. Read the study

Lumis SDKExperimental · 0.1.0

A foundation you can inspect.

A proof of concept of our investigation core, open source and read-only. Deterministic checks, one bounded investigator, and a mechanical check of every claim — on your own Kubernetes, Prometheus, Loki, Tempo, Prefect, SQL and Git.

# Checks run first. The model proposes; evidence decides.Hypothesis(    statement="v1.7 increased database reads",    causal_path=["deploy:v1.7", "feature-service", "postgres"],    predictions=["reads rose after the deployment"],    falsifiers=["amplification predates v1.7"],)
Experimental · 0.1.0 · the API may change

04 / Research directions

Different systems. A shared need to understand.

Lumis starts with data and AI systems. The rest are questions we are studying, not products we are building — each must earn its place with evidence.

Photography: Unsplash. Footage: Mixkit. Illustrative — Lumis does not operate the facilities shown.Products and directions

Follow the research.

Read how Lumis did on fifteen injected failures, try the open-source SDK, or tell us about a failure worth studying.