arXiv:2506.15065cs.LGcs.RO2025-06EMNLP被引 12

研究大模型驱动的智能体幻觉问题,发现其在复杂任务中易因环境不一致而犯错。

HEAL: An Empirical Study on Hallucinations in Embodied Agents Driven by Large Language Models

  • 构建新测试集,使幻觉率最高提升40倍
  • 12个模型在任务与场景不符时均无法正确推理
  • 揭示幻觉成因,为更可靠规划提供指导

大型语言模型(LLMs)正被广泛用作具身智能体的认知核心。然而,由于无法将用户指令与实际物理环境对齐,模型会引发幻觉,导致导航错误,如寻找不存在的冰箱。本文首次系统研究基于LLM的具身智能体在场景-任务不一致条件下执行长时程任务时的幻觉现象。通过改进现有基准,构建可诱发幻觉率高达基线40倍的探测数据集,评估了12个模型在两个仿真环境中的表现。结果表明,尽管模型具备一定推理能力,却无法解决场景与任务间的矛盾,暴露出处理不可行任务的根本缺陷。研究还提出各类场景下理想模型行为的可行建议,为开发更鲁棒、可靠的规划策略提供依据。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly being adopted as the cognitive core of embodied agents. However, inherited hallucinations, which stem from failures to ground user instructions in the observed physical environment, can lead to navigation errors, such as searching for a refrigerator that does not exist. In this paper, we present the first systematic study of hallucinations in LLM-based embodied agents performing long-horizon tasks under scene-task inconsistencies. Our goal is to understand to what extent hallucinations occur, what types of inconsistencies trigger them, and how current models respond. To achieve these goals, we construct a hallucination probing set by building on an existing benchmark, capable of inducing hallucination rates up to 40x higher than base prompts. Evaluating 12 models across two simulation environments, we find that while models exhibit reasoning, they fail to resolve scene-task inconsistencies-highlighting fundamental limitations in handling infeasible tasks. We also provide actionable insights on ideal model behavior for each scenario, offering guidance for developing more robust and reliable planning strategies.

具身智能幻觉检测大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。