I learned the most about AI not from building a model, but from operating models after they reached production. As an MLOps engineer on energy-forecasting systems for EnBW, through Ishango.ai, a failed workflow was usually detected within minutes. Understanding why it failed could take hours or sometimes almost a full day. ☕
The answer was never in one place. It was spread across the orchestrator, the monitoring dashboards, the experiment tracker, the infrastructure code, a recent deployment, a vendor feed, and a thread in Teams where someone half-remembered the last time it happened. Prefect, Datadog, MLflow, Terraform, AWS, Kafka, Snowflake: each told the truth about its own corner. None of them told me what had actually gone wrong.
What the chat window could not do
So I did what everyone does now 😂, paste the errors into a Codex, Claude Code, Cursor (coding agents) you name it😅 with some context, point them to the codebase and ask them to assist investigate the issue. The answers were fluent and often useful. They were also, too often, guesses: plausible explanations of a system the model had never seen running.
The work did not disappear; it moved. I became the model’s hands. I need access to this dashboard. Now that log. Did that deploy go out before or after the alert? I forgot to mention the vendor. Every useful answer depended on evidence I had to find, fetch, and paste — the very work I hoped to hand off. Progress, of a sort 🙃.
This is not a complaint about language models. Microsoft researchers found that large models produced root-cause and mitigation suggestions on-call engineers rated useful, across more than 40,000 real cloud incidents (Ahmed et al., ICSE 2023). The models are capable. What they lack is the system. In our own GridCast study, the same model that found the right cause 25 times in 28 inside the Lumis loop found it once in 30 when it was given only the alert.
A model that has never seen your system can only guess about it. Fluency is not evidence.
From an AI SRE to a layer
The first version of Lumis looked a lot like what people now call an AI SRE: take the incident, read the logs, explain what broke. It is a real category now. Datadog launched Bits AI SRE, an agent that investigates alerts with telemetry and architecture context, and startups like Resolve AI are building agents for production incidents. Serious teams, doing serious work. Good — it means the problem is real, and I am not imagining the pager. 📟
Building it, my question changed. Not “can a model explain this incident?” — sometimes it can — but “what would it take for this to work at scale, on estates it has never seen, with every claim checkable, and to get better with each incident instead of starting from zero?” The answer did not look like a better prompt. It looked like a missing layer.
The layer after the model
The way I think about it, language models have given software its first layer of intelligence: they read, write, reason about, and generate language and code, and they keep getting better. The layer on top of them is still mostly missing. It is the layer that knows a specific system — what depends on what, what changed, what evidence exists, what it may touch — checks before it guesses, and learns from every incident without losing human control.
Beyond that is a third layer, where software stops only describing the world and starts acting in it. Researchers call these cyber-physical systems: “interacting digital, analog, physical, and human components engineered for function through integrated physics and logic” (NIST). Energy grids, factories, robots, and vehicles. At CES 2025, NVIDIA’s Jensen Huang described the arc as perception AI, then generative AI, then physical AI: AI that can perceive, reason, plan, and act (NVIDIA). I believe that is where the next decade goes. I also believe a model that cannot yet explain a late data pipeline should not be trusted with a production line.
Somewhere to land
One paper that shaped my thinking is Tom Zahavy’s position paper “LLMs can’t jump” (Google DeepMind, ICML 2026). As I read it: generative AI has mastered induction — learning patterns from data — and is getting good at deduction — reasoning from premises it is given. What it lacks is abduction, “the generation of novel explanatory hypotheses”. Einstein described discovery as a jump from experience to new axioms, followed by deduction; the paper argues models can do the second part but not the jump, and that grounding them in physically consistent world models is the way across.
Debugging a production incident is abduction in miniature 😅. An alert is a surprising result; the root cause is the explanation nobody has written down yet. That is exactly where I watched models struggle: fluent with patterns they had seen before, shaky when the answer depended on this system, today.
Lumis does not claim to teach a model to jump. It changes what the model is asked to do. The model proposes competing explanations, each with predictions and the evidence that would prove it wrong; the system supplies the grounding — a model of the estate, registered queries, change records — and checks every claim against what it finds. If you squint, the operational graph is a small and very boring world model 🙂. Whether or not models can jump, they need somewhere to land.
Not competing with the labs
The last thing we would ever think of is to out-train Anthropic, OpenAI, Google,xAI or Moonshot AI etc.., these labs are extreme outliers are doing incredible things, and we do not need to. Training a single frontier model was estimated at roughly $78–100 million for GPT-4 and about $192 million for Gemini 1.0 Ultra, with costs growing two- to three-fold a year (Stanford AI Index 2025, with Epoch AI). Competing on that layer would be a strange use of a small team. Building on it is not.
So Lumis is deliberately model-agnostic. Today the open-source SDK can use models through OpenRouter, OpenAI, Anthropic, or Gemini, and the structure around the model — not the model itself — does the work that matters: deciding what to look at, checking what it finds, and refusing to conclude when the evidence is not there. Every time the labs ship a better model, that layer gets better for free.
What the next layer needs
If the second layer is going to be trusted with real systems, I think it needs a few properties that a coding agent chat window does not have. Some of these exist in the open-source SDK today; the rest are what we are researching.
- A model of the system. Services, data, dependencies, and recent changes, built from what the estate already reports. In the SDK today.
- Checks before guesses. Known failures matched deterministically, with no model call. The model is for what the checks cannot explain. In the SDK today.
- Bounded access. You decide where it may look: registered queries, allowlisted code, read-only sources. Context is redacted before it reaches a model — heuristically, so it is a safeguard, not a guarantee. In the SDK today.
- Your choice of model. Pick the provider you trust. Running your own hosted models is something we want to support next.
- Evidence for every claim. A conclusion is reported only when the evidence collected supports it; otherwise the answer is “insufficient evidence”. In the SDK today.
- Memory and runbooks. What was learned last time, and the procedures your team already trusts, used in the next investigation. Research.
- People in control. It reports; a person decides. Any step towards acting is earned with evidence, one boundary at a time. The principle behind all of it.
Where it came from
The idea took shape at EnBW, where I worked with Sharleen Muoki, who shared the frustration and the conviction that it could be better. It became a paper, Agentic Self-Healing for Data & AI Pipelines, written with Sharleen and our co-authors Dennis Murage, Chih-Chun Chen, Stephen Adjignon, Matteo Staar, and Oliver Angélil. Lumis is my attempt to build the hardest part of it: the reasoning.
Two earlier experiences pointed the same way. At MinoHealth AI Labs I built the tool integrations behind Moremi Bio, coordinating dozens of scientific tools for multi-step research workflows. It taught me that an intelligent system depends as much on how computation and evidence are organised as on the model. At Noeud, combining probabilistic forecasts with market data and bounded LLM reasoning raised a related question: how do you let a generative model use evidence without letting it overwrite a more reliable source? Noeud is also a natural place to try Lumis outside a synthetic estate — when it is ready, not before.
Why start with data and AI
Because I have lived it. Data and ML pipelines fail in ways that cross every boundary — code, data, infrastructure, vendors, models — and the cost is real: in ITIC’s 2024 survey, more than 90% of mid-size and large enterprises put the cost of a single hour of downtime above $300,000 (ITIC). These failures are also reproducible, which makes them honest to study. You can inject one, freeze the evidence, and check whether a system found the cause or just sounded like it had. That is what GridCast is for.
Where it goes
If the second layer works for data and AI systems, the same idea should carry over to cloud infrastructure, then to edge devices, and eventually to cyber-physical systems, where software, sensors, machines, and people meet. That is a research direction we will be taking on in the coming years. The bar rises with every step, because the cost of a wrong action rises with it. One day that may mean working alongside labs that build robots and physical systems; it will never mean skipping the evidence.
Control rooms once made the state of a power plant legible to the people running it. I want software to do the same for systems too complex to read at a glance — and to stay accountable while it does. Language models gave us the first layer. Lumis is my bet on the next one.
If this is a problem you think about too — as a researcher, an engineer, or someone who has been paged about a pipeline at 2 a.m. — I would genuinely like to compare notes. Say hello. 👋
