arXiv:2605.24219cs.AI2026-05

检测多智能体工业流程中的中间步骤幻觉,发现近半数错误涉及多种幻觉类型。

Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows

论文配图:Beyond Final Answers: Auditing Trajectory-Level Hallucinations in Multi-Agent Industrial Workflows
图 1 · 摘自论文原文
  • 构建五类幻觉分类体系,基于专家标注的智能体轨迹进行评估
  • 近半数幻觉轨迹同时包含多种错误类型,现有基准遗漏关键问题
  • 轨迹感知检测显著优于事后验证,适合部署前安全审计

大型语言模型正被用作能推理、调用工具并执行多步任务的自主智能体。然而,现有幻觉评估仍仅关注最终输出,忽略了中间思维-行动-观察步骤中的错误。本文提出Trajel数据集与评估框架,用于审计多智能体工业工作流中的轨迹级幻觉。Trajel基于AssetOpsBench的专家标注轨迹,引入五类幻觉分类(事实性、指代性、逻辑性、程序性、范围性)。我们在子任务、轨迹和长上下文三个层面测试监督式检测模型。结果显示,最常见的错误类型被现有基准忽略;近一半幻觉轨迹同时包含多种类型;即使自动化检测器在二分类上准确率高,仍会误判最细微的类型。轨迹感知检测显著优于传统事后验证,表明基于分类体系的评估对更安全的智能体部署至关重要。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed as autonomous agents that reason, use tools, and act over multiple steps. Yet most hallucination benchmarks still evaluate only the final output, missing failures that originate in intermediate Thought-Action-Observation steps. We present Trajel, a dataset and evaluation framework for auditing trajectory-level hallucinations in multi-agent industrial workflows. Trajel introduces a five-type hallucination taxonomy (factual, referential, logical, procedural, and scope-based) over expert-annotated agent traces from AssetOpsBench. We benchmark supervised detection models at the subtask, trajectory, and long-context levels. Our results show that the most common failure modes are missed by existing benchmarks, that nearly half of hallucinated trajectories involve multiple types at once, and that automated detectors with high binary accuracy still misclassify the subtlest types. Trajectory-aware detection significantly outperforms standard post-hoc verification, making taxonomy-grounded evaluation necessary for safer agentic deployment.

幻觉检测多智能体工业流程LLM评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。