用开源工具构建可自愈的数据与AI管道架构,省钱省事。
Agentic Self-Healing for Data and AI Pipelines: An Affordable Vendor-Agnostic Architecture using Open-Source Software

- 整合监控、诊断、修复等能力,打造可自愈的管道系统。
- 基于开源组件实现跨平台兼容,避免厂商锁定。
- 适合中小型团队快速搭建低投入的智能运维体系。
现代组织依赖数据、机器学习和软件交付流水线来处理数据、训练模型、部署应用、更新仪表盘并支持关键业务决策。然而,这些流水线常因数据质量、模式变更、上游源变化、基础设施问题、编排失败及模型流程异常而失效。现有零运维、可观测性及AI运维平台虽能检测故障、分析根因,部分还可推荐或执行修复,但多数成本高、厂商绑定强,小团队难以适配多工具环境。本文首先对比了现成的AI辅助流水线监控、根因分析与自动修复方案,发现核心短板在于架构而非技术:自愈所需组件已存在,但分散在不同厂商平台、可观测工具、告警系统与开源组件中。因此,我们提出一种低成本、厂商无绑定的参考架构,利用开源工具实现代理式自愈流水线。该架构融合监控、流水线元数据、事件历史、确定性策略检查、AI辅助诊断、审批流程与受控修复动作,助力团队以更少人工干预完成故障检测、诊断、修复、验证与复盘。目标是提供一个可广泛适配数据工程、MLOps与软件交付场景的实用架构。
原文摘要 · Abstract (English)
Modern organizations rely on data, machine learning, and software delivery pipelines to move data, train models, deploy applications, refresh dashboards, and support business-critical decisions. However, these pipelines often fail because of data quality issues, schema changes, upstream source changes, infrastructure problems, orchestration failures, and model workflow issues. Existing ZeroOps, observability, and AI operations platforms can help teams detect incidents, investigate root causes, and in some cases recommend or execute fixes. However, many of these solutions are expensive, vendor-specific, or difficult for smaller teams to adapt across different tools and environments. This paper first compares existing off-the-shelf solutions for AI-assisted pipeline monitoring, root-cause analysis, and automated remediation, including their strengths, limitations, and practical trade-offs. Based on this comparison, we find that the main gap is architectural rather than technological: the required ingredients for self-healing pipelines already exist, but they are fragmented across vendor-specific platforms, observability tools, incident systems, and open-source components. We therefore propose an affordable, vendor-agnostic reference architecture for agentic self-healing pipelines using open-source and low-cost tools. The proposed architecture combines monitoring, pipeline metadata, incident history, deterministic policy checks, AI-assisted diagnosis, approval workflows, and controlled remediation actions to help teams detect, diagnose, repair, verify, and learn from pipeline issues with less manual effort. The goal is to provide a practical reference architecture that can be adapted across data engineering, machine learning operations, and software delivery environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。