arXiv:2507.21017cs.AI2025-07被引 13

首个系统评估大模型代理幻觉的基准,定位错误行为根源。

MIRAGE-Bench: LLM Agent is Hallucinating and Where to Find Them

  • 构建三类幻觉分类体系,覆盖指令、历史与环境三方面偏差。
  • 通过快照策略生成可复现测试用例,精准定位决策失误点。
  • 用定制化评分机制实现高效高保真评估,适合安全研究者使用。

幻觉对基于大语言模型(LLM)的智能体构成重大风险,常表现为因认知上下文中虚构或误读信息而产生的错误动作。尽管已有研究揭示此类问题,但现有评估分散且缺乏统一测试平台。本文提出 MIRAGE-Bench——Measuring Illusions in Risky AGEnt settings,首个用于诱发和评估交互式 LLM 代理幻觉的统一基准。我们首先引入一个三部分分类体系,以识别代理在任务指令、执行历史和环境观察三个方面出现的不忠实行为。通过系统审计现有代理基准,利用快照策略隔离确定性与可复现的决策点,合成测试用例。为评估幻觉行为,采用细粒度的 LLM-as-a-Judge 框架,配合定制的风险感知提示,实现无需枚举完整动作空间即可进行可扩展、高保真的评估。MIRAGE-Bench 为理解智能体失败模式提供切实洞见,并为交互环境中抑制幻觉奠定基础。

原文摘要 · Abstract (English)

Hallucinations pose critical risks for large language model (LLM)-based agents, often manifesting as hallucinative actions resulting from fabricated or misinterpreted information within the cognitive context. While recent studies have exposed such failures, existing evaluations remain fragmented and lack a principled testbed. In this paper, we present MIRAGE-Bench--Measuring Illusions in Risky AGEnt settings--the first unified benchmark for eliciting and evaluating hallucinations in interactive LLM-agent scenarios. We begin by introducing a three-part taxonomy to address agentic hallucinations: actions that are unfaithful to (i) task instructions, (ii) execution history, or (iii) environment observations. To analyze, we first elicit such failures by performing a systematic audit of existing agent benchmarks, then synthesize test cases using a snapshot strategy that isolates decision points in deterministic and reproducible manners. To evaluate hallucination behaviors, we adopt a fine-grained-level LLM-as-a-Judge paradigm with tailored risk-aware prompts, enabling scalable, high-fidelity assessment of agent actions without enumerating full action spaces. MIRAGE-Bench provides actionable insights on failure modes of LLM agents and lays the groundwork for principled progress in mitigating hallucinations in interactive environments.

幻觉检测智能体评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。