A dashboard goes stale at nine in the morning. Somewhere upstream, a source team renamed a column, an orchestrator retried a task four times, and a feature pipeline quietly fed yesterday’s data to today’s model. The alert tells you something is wrong. It does not tell you what, where, or whether the fix you are about to apply will actually work.
That gap — between noticing a failure and safely recovering from it — is what our paper, Agentic Self-Healing for Data & AI Pipelines, sets out to close. Its central finding is simple, and a little uncomfortable for anyone selling a platform: the technology needed for self-healing pipelines already exists. What is missing is architecture.
Where this came from
The paper did not start with a model. It started with production. I operated data and ML workflows for energy forecasting, and most of the incidents I worked on were easy to detect and hard to explain. An alert fired within minutes; understanding why could take hours, because the answer was spread across the orchestrator, the infrastructure, recent deployments, upstream data, and the monitoring stack.
Earlier work pointed the same way. Building agentic research systems that coordinated dozens of scientific tools taught me that an AI system depends as much on how its evidence and computation are organised as on the model itself. Combining forecasts with market data and bounded language-model reasoning raised a sharper version of the question: how do you let a model use many kinds of evidence without letting it overrule the more reliable ones? The paper is our first answer for pipelines.
Pipelines fail in six familiar ways
Across data engineering, machine learning operations, and software delivery, failures cluster into a small number of recurring classes. They surface in different places — the data plane, the control plane, model quality — but they share one diagnostic structure: the symptom usually appears several hops downstream of the cause.
- Data quality null spikes, duplicates, and distribution shifts that break downstream assumptions.
- Schema changes columns renamed, retyped, or dropped without notice.
- Upstream sources API versions, export cadences, revoked credentials.
- Infrastructure out-of-memory kills, preemption, expiring certificates, quotas.
- Orchestration deadlocks, stuck sensors, backfill collisions, retry storms.
- Model workflows training divergence, feature skew, slow performance decay.
Today, the recovery path is mostly human: an engineer reconstructs the failure across logs, orchestrator screens, and upstream systems, applies a fix by hand, and watches the next run. Every step leans on tacit knowledge held by a few senior people.
What eight platforms taught us
We compared eight commercial offerings for AI-assisted monitoring, root-cause analysis, and remediation — spanning platform-native autonomous operations, data observability, application observability, incident response, and IT service management. They are genuinely capable. Three patterns stood out.
- Autonomy correlates with lock-in. The platforms that close the loop most completely do so because they own the surrounding context — the catalog, the orchestrator, the CMDB. Healing is easy inside one vendor’s boundary; it is the heterogeneous estate that is hard.
- No single product covers every pipeline domain. Data observability is thin on remediation, application observability is thin on table- and model-level semantics, and incident tools depend on others for detection. Teams end up paying for two or three.
- The mechanisms are convergent. Underneath the branding, every platform combines the same five ingredients: telemetry, metadata and topology, an inference layer, a policy gate, and an action runtime. Each now has a mature open-source counterpart.
What is missing is not technology but architecture: a reference design that shows how to assemble these commodity parts into a trustworthy self-healing loop with appropriate human oversight.
A seven-layer reference architecture
The paper proposes Agentic Recovery and Incident Response: a vendor-agnostic architecture defined by responsibilities and interfaces rather than by a tool list. The pipelines being healed stay exactly as they are. Everything else is a replaceable layer.
- Existing estate orchestrators, transforms, ML workflows, delivery — unchanged, only instrumented.
- Telemetry & signals metrics, logs and traces, lineage, data-quality checks through open standards.
- Incident memory a deliberately boring store of past incidents, runbooks, and pipeline metadata.
- Policy & reasoning a deterministic policy engine first; LLM agents for triage, diagnosis, planning, and verification only when rules cannot resolve it.
- Approval & governance risk tiers decide what is auto-approved, what needs a person, and what is recommendation-only.
- Guarded execution allowlisted, rate-limited, reversible actions with an immutable audit log.
- Verification & learning re-run the failed checks, watch recovery, write the whole episode back to memory — and promote recurring fixes into deterministic rules.
Six principles that invert the trade-offs
- Vendor agnosticism through open interfaces. OpenTelemetry, Prometheus, OpenLineage — every component is replaceable.
- The LLM is a swappable commodity. A model gateway makes hosted, routed, and local models interchangeable per task.
- Deterministic before generative. Known failures go to rules and playbooks; models are reserved for ambiguous incidents.
- Guarded autonomy, not full autonomy. Agents act only from an allowlist, and every action carries a risk tier.
- Memory is the moat. Each resolved episode grounds the next diagnosis.
- Incremental adoption. Each layer is useful alone: monitor first, diagnose later, act last.
The loop, in practice
Take schema drift. An upstream team renames customer_id to cust_id. A lightweight validation check fails at ingestion; triage classifies the change and uses lineage to flag the downstream models and dashboards at risk. Diagnosis compares the schema with the last good run, finds the rename, and retrieves a similar incident from months earlier. The planner selects an approved playbook that opens a pull request — a reviewable action, so one engineer approves. Verification re-runs the checks and confirms freshness downstream, and the episode is written back to memory.
The goal is not maximal autonomy. Agents compress investigation and execution mechanics; irreversible, high-risk, or business-sensitive decisions stay with people. Even when no playbook applies, the on-call engineer receives a completed investigation instead of a bare alert.
Where we are honest about the limits
The architecture is not an argument against buying. A team already standardised on one platform may get most of this by switching it on. The components are free to license, not free to run: telemetry, playbooks, and thresholds are real engineering work. And LLM-driven operations carry specific risks — confabulated diagnoses, approvals that decay into rubber stamps, and remediations that trigger further remediations. Grounding in evidence, rate and blast-radius limits, and reversible mechanisms exist precisely to bound them.
From reference architecture to Lumis
The paper is where Lumis began. It describes how to assemble a guarded self-healing loop; Lumis is our work on the hardest part of that loop — reasoning. The questions we care about most are the ones the paper leaves open: how a system forms competing explanations from logs, graphs, and changes; how it finds the one observation that would separate them; and when it should act, abstain, or hand the decision to a person. In Lumis, rules, memory, and language models can all propose explanations, but only evidence decides which survives. We have since tested that reasoning against GridCast, a reproducible reference estate with fifteen injected failures, and published what we found — including what went wrong.
Cite the paper
Solomon Eshun, Dennis Murage, Sharleen Muoki, Chih-Chun Chen, Stephen Adjignon, Matteo Staar, Oliver Angélil. Agentic Self-Healing for Data & AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software. arXiv:2608.01955, 2026.
@misc{eshun2026agentic,
title = {Agentic Self-Healing for Data \& AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software},
author = {Eshun, Solomon and Murage, Dennis and Muoki, Sharleen and Chen, Chih-Chun and Adjignon, Stephen and Staar, Matteo and Ang{\'e}lil, Oliver},
year = {2026},
eprint = {2608.01955},
archivePrefix = {arXiv},
primaryClass = {cs.ET}
}