arXiv:2607.00828cs.DBcs.AI2026-07被引 1

研究智能体生成分析流程时的语义断层问题,发现153次失败源于概念与数据的不匹配。

Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows

论文配图:Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows
图 1 · 摘自论文原文
  • 通过跨领域实证研究,识别出五类分析意图无法落地的核心原因。
  • 236个分析需求中153次失败,揭示用户概念与数据表达间的深层语义鸿沟。
  • 适合关注AI数据分析系统可靠性、提示工程与知识表示的研究者参考。

大型语言模型(LLMs)被广泛用于生成查询、调用工具并构建分析工作流。尽管近期进展显著提升了工作流的生成与执行能力,但将分析概念转化为可操作计算所需的关键语义信息,往往超出数据库模式和数据值的显式表达范围。我们开展了一项跨领域的形成性研究,分析智能体生成的工作流在实际操作中的失败情况。在涵盖金融、人力资源和公共安全领域的236个分析意图中,尽管工作流生成与执行成功,仍出现153次操作化失败。分析揭示出五类重复出现的失败类型:比较基准缺失、过程推理偏差、量化推理错误、角色混淆及政策语义未对齐。这些发现表明,用户层面的分析概念与工作流生成系统可获取的信息之间存在显著语义断层。更广泛地,这引发了对分析操作合理性的质疑,暗示未来的智能体数据系统可能需要更丰富的语义表征,以弥合分析意图与可执行计算之间的差距。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly used to generate queries, invoke tools, and construct analytical workflows. Although recent advances have substantially improved workflow generation and execution, the semantic information required to operationalize analytical concepts often lies beyond what is explicitly represented in database schemas and data values. We present a cross-domain formative study of operationalization failures in agent-generated analytical workflows. Across 236 analytical intents spanning finance, human resources, and public safety domains, we identify 153 recurring failures despite successful workflow generation and execution. Our analysis reveals five recurring classes of failures: comparative grounding, process reasoning, quantitative reasoning, role confusion, and policy grounding. These findings suggest a semantic gap between user-level analytical concepts and the information available to workflow-generation systems. More broadly, they raise questions about the admissibility of analytical operations and suggest that future agentic data systems may require richer semantic representations to bridge the gap between analytical intent and executable computation.

智能体系统语义鸿沟数据分析大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。