arXiv:2510.09567cs.AIcs.DB2025-10被引 10

让不信任的AI代理安全操作数据湖仓,通过代码证明机制保障可靠性。

Safe, Untrusted, "Proof-Carrying" AI Agents: toward the agentic lakehouse

  • 用声明式环境与数据分支支持可复现的智能体工作流
  • 原型验证了代理在生产数据上修复管道时的正确性与安全性
  • 适合关注数据治理与自动化安全的工程团队

数据湖仓承载敏感任务,基于AI的自动化引发对信任、正确性和治理的担忧。我们提出,以API优先、可编程的数据湖仓为安全设计的智能体工作流提供了合适抽象。以Bauplan为例,展示数据分支与声明式环境如何自然扩展至智能体,实现可复现性与可观测性,同时缩小攻击面。我们构建了一个原型,利用受代码证明启发的正确性检查,使智能体能够修复数据管道。结果表明,未经信任的AI代理可在生产数据上安全运行,并为全智能体湖仓铺平道路。

原文摘要 · Abstract (English)

Data lakehouses run sensitive workloads, where AI-driven automation raises concerns about trust, correctness, and governance. We argue that API-first, programmable lakehouses provide the right abstractions for safe-by-design, agentic workflows. Using Bauplan as a case study, we show how data branching and declarative environments extend naturally to agents, enabling reproducibility and observability while reducing the attack surface. We present a proof-of-concept in which agents repair data pipelines using correctness checks inspired by proof-carrying code. Our prototype demonstrates that untrusted AI agents can operate safely on production data and outlines a path toward a fully agentic lakehouse.

智能体数据湖仓安全自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。