arXiv:2605.01428cs.CL2026-05被引 3

让大模型学会承认不确定,比盲目自信更可信。

Hallucinations Undermine Trust; Metacognition is a Way Forward

论文配图:Hallucinations Undermine Trust; Metacognition is a Way Forward
图 1 · 摘自论文原文
  • 把幻觉视为自信的错误,主张用真实不确定性替代盲目回答
  • 模型在简单问答中仍频繁出错,因缺乏对知识边界的自我认知
  • 适合追求可信AI的开发者与系统设计者

尽管生成式AI在事实准确性上取得进展,但错误(常称幻觉)仍是核心问题,尤其在复杂或细微场景中。即使在最简单的事实型问答任务中,无外部工具的前沿模型仍会幻觉。我们指出,当前准确性的提升主要来自扩大知识边界(编码更多事实),而非增强对边界的认知能力(区分已知与未知)。我们认为后者的难度本质在于:模型可能无法完美区分真伪,导致消除幻觉与保持实用性之间存在不可避免的权衡。这一权衡在新视角下可化解——若将幻觉定义为未经恰当说明的自信错误,便出现第三条路径:表达不确定性。我们提出‘真实不确定性’,即语言上的不确定与内在不确定相一致。这是元认知的一方面:意识到自身不确定并据此行动。对于直接交互,意味着诚实表达不确定;对于代理系统,则成为控制何时搜索、信任什么的底层机制。因此,元认知对大模型既可信又高效至关重要。最后,我们指出了通往该目标的开放性问题。

原文摘要 · Abstract (English)

Despite significant strides in factual reliability, errors -- often termed hallucinations -- remain a major concern for generative AI, especially as LLMs are increasingly expected to be helpful in more complex or nuanced setups. Yet even in the simplest setting -- factoid question-answering with clear ground truth-frontier models without external tools continue to hallucinate. We argue that most factuality gains in this domain have come from expanding the model's knowledge boundary (encoding more facts) rather than improving awareness of that boundary (distinguishing known from unknown). We conjecture that the latter is inherently difficult: models may lack the discriminative power to perfectly separate truths from errors, creating an unavoidable tradeoff between eliminating hallucinations and preserving utility. This tradeoff dissolves under a different framing. If we understand hallucinations as confident errors -- incorrect information delivered without appropriate qualification -- a third path emerges beyond the answer-or-abstain dichotomy: expressing uncertainty. We propose faithful uncertainty: aligning linguistic uncertainty with intrinsic uncertainty. This is one facet of metacognition -- the ability to be aware of one's own uncertainty and to act on it. For direct interaction, acting on uncertainty means communicating it honestly; for agentic systems, it becomes the control layer governing when to search and what to trust. Metacognition is thus essential for LLMs to be both trustworthy and capable; we conclude by highlighting open problems for progress towards this objective.

元认知幻觉可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。