Network equipment in a dark data-centre rack

Lumis Data & AI · R&D · research prototype

Pipeline failures, explained with evidence, not guesses.

Diagnosis for data pipelines, ML workflows, and the infrastructure beneath them. Deterministic checks first. A bounded investigator when they fall short. A person in control.

Photograph: Tyler / Unsplash

What it understands

The vocabulary of data and ML systems.

Lumis Data & AI models your estate in the terms your engineers already use — so an incident is a broken relationship between real things, not a wall of alerts.

Entities & signals
DatasetPipelineModelFeatureWorkflowSchemaLineageFreshnessDriftDeployment
Works today · read-only
KubernetesPrometheusLokiTempoPrefectPostgreSQLGit

Connectors describe what Lumis may read. On the roadmap: Airflow, Databricks, Snowflake, MLflow, Kafka, OpenLineage.

How a diagnosis happens

Checks first. An investigator only when needed.

Known failures can be matched by checks, with no model call. Only what the checks cannot explain goes to a single bounded investigator, with a budget and read-only tools. Below are three real incidents from the GridCast evaluation, as Lumis reported them.

  1. 01
    Normalise and scope
    alert · PlanningApiUnreachableplanning-api · replicas wanted 0planning-api · replicas available 0grid-operator · failed plan reads ↑
  2. 02
    Deterministic checks · no model call
    • planning-api-scaled-to-zeroMATCHTerminal: the signature explains the alert on its own
    • feature-query-amplificationNO MATCH
    • forecast-service-oom-killedNO MATCH
    • forecast-model-slowdownNO MATCH
  3. 03
    The check explains it

    Resolved by the check alone in under a second: 14 evidence queries, no model call. The cause: planning-api had been left at zero replicas after a maintenance window.

  4. 04
    Guardrails

    Every claim is checked against the evidence collected. If it does not hold: “insufficient evidence”.

  5. 05
    Human review

    A structured report reaches an engineer. Conclusion: supported_diagnosis. Correct in all 7 runs: 5 rules-only, 2 Lumis.

Real runs from the GridCast evaluation (4 October 2026; one synthetic estate, one model).Scenario J report and transcript The full study

What you get back

A report you can audit, not a paragraph to trust.

Every investigation ends in structure: the explanations considered, the evidence for and against each, and what is still unknown. Auditable artefacts are stored — not a model's stream of thought.

Investigation report · GridCast scenario A · real runsupported_diagnosis

Forecast pipeline above 5 s after feature-service 1.7.0

  • H1Release 1.7.0 switched lag features to minute resolutionsupported
  • checkplanning-api scaled to zeroruled out
  • checkFeature builds or database auth failingruled out
  • checkforecast-service OOM or model slowdownruled out
Evidence
30 queries · 6 cited
Route
checks → investigator
Unresolved
Release flag inferred from the release catalogue, not read from the pod
Possible conclusions: supported_diagnosis · insufficient_evidence · requires_human_expert. Read the full report

Follow the research.

Lumis Data & AI is a research prototype. Read how it performed on fifteen injected failures, try the open-source SDK, or tell us about a failure class you would like studied.