新指标衡量语言模型幻觉,区分盲目自信与真实不确定。
Measuring Language Model Hallucinations Through Distributional Correctness
- 用概率分布评估模型回答,而非仅看对错
- 一半测试基准下所有模型得分均为负,显示严重幻觉
- 适合关注模型可信度与安全性的研究者
现有语言模型评估多依赖单次回答的准确率或评分规则,未能捕捉模型信念状态的全貌。近期研究指出,模型幻觉部分源于其在二元评分机制下被优化为‘答题’而非‘放弃’。尽管已有惩罚方法,却忽略了模型在不确定性表达上的差异——例如倾向于错误答案还是‘我不知道’。为此提出分布正确性评分(DCS),可衡量模型在所有可能答案上的概率分布,自然区分有害的错误自信与通过弃权表达的真实不确定。理论分析与实例表明,DCS提供更精细、更契合目标的评估方式,激励模型真实表达不确定。将12个现有评测基准适配至DCS变体,并在6个语言模型上测试发现,半数基准下所有模型得分均为负,表明普遍存在显著幻觉倾向。
原文摘要 · Abstract (English)
Common evaluation paradigms for language models focus on scoring single responses through accuracy metrics or proper scoring rules, failing to capture the full richness of a model's belief state. Recent work illustrates that language models hallucinate in-part because they are optimised to be good test-takers under binary scoring schemes that reward any answer over abstention. While this insight naturally leads to penalty-based approaches, they ignore crucial distinctions in how models distribute uncertainty, for example between hedging toward incorrect answers versus hedging toward "I don't know" responses. A novel evaluation metric, the Distributional Correctness Score (DCS), is introduced to solve this problem, i.e., of not considering a model's entire probability distribution over answer choices. DCS naturally distinguishes between harmful overconfidence in wrong answers and uncertainty expressed through abstention, providing scores in an interpretable default range. Through theoretical analysis and illustrative examples, DCS is demonstrated to offer a more nuanced and aligned evaluation paradigm that incentivises models to express genuine uncertainty rather than guessing. Adapting 12 existing evaluation benchmarks to DCS's variants and measuring performance on six language models reveals that for half of the tested benchmarks scores are negative across all tested models, indicating significant tendencies towards hallucination.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。