All projects View code
Agentic Data Reliability
Self-Healing Data Pipeline Agent
An agentic ETL system that detects schema drift, drafts controlled fixes, requires human approval, validates recovery, and automatically rolls back failed mutations.
THE PROBLEM
Why this exists.
Data pipelines fail when upstream schemas drift. Manual triage is repetitive, slow, and risky when automated fixes are allowed to mutate production data.
THE APPROACH
How I approached it.
The agent classifies drift, drafts a transformation, gates the change behind approval, executes it in an isolated path, re-validates the schema, and restores the last known-good state if validation fails.
ARCHITECTURE
The system flow.
Incoming batch
Drift detector
Fix drafter
Human approval
Sandboxed executor
Validation
Audit / rollback
ENGINEERING SIGNALS
What to notice.
- 4 schema-drift classes
- Human-in-the-loop approval
- Automatic rollback
- Structured audit trail
- 130+ pytest tests
WHAT I LEARNED
The takeaway.
A useful autonomous system needs boundaries. The most important engineering decision was not “how do I let an agent fix data?” but “how do I make every proposed change observable, reviewable, reversible, and testable?”