arXiv:2606.19319cs.MAcs.AI2026-06

用自主编程代理自动处理企业数据的解读、建模与查询。

Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents

论文配图:Data Intelligence Agents: Interpreting, Modeling, and Querying Enterprise Data via Autonomous Coding Agents
图 1 · 摘自论文原文
  • 三个代理协同工作,通过执行代码生成并验证数据成果。
  • 在七个SQL基准测试中表现优于或持平现有最佳结果。
  • 适合需要自动化数据处理的企业用户和数据工程师。

生产环境中的数据集成因数据所有者、工程师和分析师之间的重复、低效协作而受阻,需共同完成数据发现、结构化和查询。我们提出数据智能代理(DIA),由三个代理(数据解释器、模式创建器、查询生成器)构成,将自主编程代理(ACAs)作为核心抽象:代理不输出文本,而是生成、执行、验证和修复具体成果,共享记忆以复用经验,并将结果提交给领域专家审查。DIA已在企业客户中部署。我们深入研究了查询生成器,在完全自治模式下评估其在涵盖四个任务类别和四种方言的七个SQL基准上的表现。其在全部七项测试中达到或超过已有最优结果,证明基于执行、依托自主编程代理和共享记忆的架构,能有效泛化于数据智能任务,仅需自然语言指令即可适应不同场景。

原文摘要 · Abstract (English)

Production data integration is bottlenecked by repeated, lossy handoffs between data owners, engineers, and analysts who must collaboratively discover, structure, and query enterprise data. We present Data Intelligence Agents (DIA), a system of three agents (Data Interpreter, Schema Creator, and Query Generator) that compresses this workflow by treating autonomous coding agents (ACAs) as a first-class abstraction: rather than emitting text, the agents generate, execute, validate, and repair concrete artifacts, draw on a shared memory for experience reuse, and surface each for review by domain experts. DIA is deployed in production for enterprise customers. We study the Query Generator in depth and evaluate it in fully autonomous mode across seven SQL benchmarks spanning four task categories and four dialects. It matches or surpasses the best published results on all seven, demonstrating that an architecture grounded in execution, built on ACAs and a shared memory, generalizes across the data intelligence workload with adaptation confined to natural-language instructions.

数据智能自主编程企业数据SQL生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。