构建可控制的参考世界,系统评估大模型幻觉成因与表现。
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models

- 以明确参考世界为基准,自动标注幻觉行为。
- 前沿模型在感知幻觉上已接近解决,但多步推理仍困难。
- 适合研究幻觉机制或评估模型可信度的研究者使用。
幻觉仍是大语言模型的核心缺陷,现有基准在摘要、问答、检索增强生成和代理交互等任务中对幻觉的定义不一致,导致难以判断某种缓解方法是否跨场景有效。当前基准要么依赖人工标注且参考内容固定易被记忆,要么基于难以复现的观测场景。为探究根本原因,我们提出HalluWorld,一种基于显式参考世界范式的可扩展基准:当模型生成与该世界不符的可观测陈述时即视为幻觉。在此框架下,我们构建了合成与半合成环境,其中参考世界完全指定,模型视角可控,幻觉标签可自动生成。HalluWorld涵盖网格世界、国际象棋和真实终端任务,支持对世界复杂度、可观测性、时间变化和源冲突策略的可控变化,并将幻觉细分为多个错误类别。我们在这些场景中评估了前沿及开源语言模型,发现:对于直接观测信息的感知幻觉,前沿模型已基本解决;而多步状态跟踪和因果前向模拟仍具挑战,且扩展思维并未普遍缓解此问题。在终端任务中,模型也难以判断何时应放弃回答。不同探测类型和领域间失败模式的不均衡表明,幻觉源于多种独立失效机制而非单一能力缺陷。结果表明,可控参考世界为测量和减少现代语言模型的幻觉提供了可扩展且可复现的路径。
原文摘要 · Abstract (English)
Hallucination remains a central failure mode of large language models, but existing benchmarks operationalize it inconsistently across summarization, question answering, retrieval-augmented generation, and agentic interaction. This fragmentation makes it unclear whether a mitigation that works in one setting reduces hallucinations across contexts. Current benchmarks either require human annotation and fixed references that may be memorized, or rely on observations in settings that are difficult to reproduce. To study root causes, we introduce HalluWorld, an extensible benchmark grounded in an explicit reference-world formulation: a model hallucinates when it produces an observable claim that is false with respect to this world. Building on this view, we construct synthetic and semi-synthetic environments in which the reference world is fully specified, the model's view is controlled, and hallucination labels are generated automatically. HalluWorld spans gridworlds, chess, and realistic terminal tasks, enabling controlled variation of world complexity, observability, temporal change, and source-conflict policy, and disentangling hallucinations into fine-grained error categories. We evaluate frontier and open-weight language models across these settings and find consistent patterns: perceptual hallucination on directly observed information is near-solved for frontier models, while multi-step state tracking and causal forward simulation remain difficult and are not generally solved by extended thinking. In the terminal setting, models also struggle with when to abstain. The uneven profile of failures across probe types and domains suggests that hallucinations arise from distinct failure modes rather than a single capability. Our results suggest that controlled reference worlds offer a scalable and reproducible path toward measuring and reducing hallucinations in modern language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。