arXiv:2511.16402cs.AIcs.DB2025-11中稿 · the Trustworthy Ag…被引 6

为智能体设计可信数据湖仓,解决生产级信任难题

Trustworthy AI in the Agentic Lakehouse: from Concurrency to Governance

  • 以事务为中心重构湖仓,支持多语言智能体并发访问
  • 提出Bauplan架构,实现数据与计算的隔离保障
  • 可自愈流水线示范,兼顾推理与可信治理

尽管人工智能能力不断提升,多数企业仍认为智能体不够可信,无法处理生产数据。本文认为,构建可信智能体工作流的关键在于先解决基础设施问题:传统湖仓不适应智能体访问模式。若以事务为核心重新设计,则治理机制自然形成。我们类比数据库中的多版本并发控制(MVCC),揭示其在解耦、多语言环境下的移植失败原因。随后提出面向智能体的架构设计Bauplan,重新实现湖仓中的数据与计算隔离。最后展示一个参考实现:自愈流水线,无缝融合智能体推理与正确性、可信性等各项保障。

原文摘要 · Abstract (English)

Even as AI capabilities improve, most enterprises do not consider agents trustworthy enough to work on production data. In this paper, we argue that the path to trustworthy agentic workflows begins with solving the infrastructure problem first: traditional lakehouses are not suited for agent access patterns, but if we design one around transactions, governance follows. In particular, we draw an operational analogy to MVCC in databases and show why a direct transplant fails in a decoupled, multi-language setting. We then propose an agent-first design, Bauplan, that reimplements data and compute isolation in the lakehouse. We conclude by sharing a reference implementation of a self-healing pipeline in Bauplan, which seamlessly couples agent reasoning with all the desired guarantees for correctness and trust.

智能体湖仓可信AI数据治理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。