arXiv:2512.21577cs.CLcs.AI2025-12被引 6

统一定义幻觉:本质是模型对世界认知的错误。

A Unified Definition of Hallucination: It's The World Model, Stupid!

  • 将幻觉定义为内部世界模型的不准确,可被用户察觉
  • 通过参考世界模型和冲突策略实现不同定义的统一
  • 适合评估、对比与改进大模型幻觉问题的研究者

尽管自语言模型诞生以来已多次尝试缓解幻觉问题,前沿大模型中仍存在持续性的幻觉现象。我们回顾了现有幻觉定义,并将其整合为一个统一框架:幻觉本质上是模型在可被用户观察到的范围内,对内部世界模型的不准确建模。例如,陈述与知识库矛盾的事实,或生成与原文相悖的摘要。通过调整参考世界模型和冲突策略,该框架可涵盖以往各类定义。这一统一视角有助于明确评估中的参考‘世界’假设,区分真实幻觉与规划或奖励错误,并为跨基准评估和治理策略讨论提供通用语言。基于此,我们还将框架与HalluWorld——一个提供完整参考世界模型以压力测试幻觉的基准——联系起来。

原文摘要 · Abstract (English)

Despite numerous attempts at mitigation since the inception of language models, hallucinations remain a persistent problem even in today's frontier LLMs. Why is this? We review existing definitions of hallucination and fold them into a single, unified definition wherein prior definitions are subsumed. We argue that hallucination can be unified by defining it as simply inaccurate (internal) world modeling, in a form where it is observable to the user. For example, stating a fact which contradicts a knowledge base OR producing a summary which contradicts the source. By varying the reference world model and conflict policy, our framework unifies prior definitions. We argue that this unified view is useful because it forces evaluations to clarify their assumed reference "world", distinguishes true hallucinations from planning or reward errors, and provides a common language for comparison across benchmarks and discussion of mitigation strategies. Building on this definition, we also connect our framework to HalluWorld, a complementary benchmark that instantiates fully specified reference world models for stress-testing model hallucinations.

幻觉定义世界模型评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。