提出新评估框架,揭示大模型幻觉中一致性问题比正确性更关键
Rethinking Hallucinations: Correctness, Consistency, and Prompt Multiplicity
- 引入提示多样性框架,量化模型输出的一致性
- 发现医学基准中超过50%的输出不一致,远超正确性问题
- 适合关注幻觉检测与模型可信度的研究者
大型语言模型常生成虚假或误导性内容,即“幻觉”,带来信任危机与信息泛滥。现有评估仅关注正确性,忽略一致性这一关键维度。本文提出提示多样性框架,用于量化模型输出一致性。分析显示,在Med-HALT等基准上,不一致性超过50%,表明幻觉危害被严重低估。进一步研究发现:(a) 检测技术实际识别的是不一致性而非正确性错误;(b) RAG等缓解方法虽有帮助,却可能引入新不一致性。通过整合提示多样性,本文构建了更全面的幻觉危害评估框架,并揭示当前检测与缓解策略的关键局限。
原文摘要 · Abstract (English)
Large language models (LLMs) are known to "hallucinate" by generating false or misleading outputs. Hallucinations pose various harms, from erosion of trust to widespread misinformation. Existing hallucination evaluation, however, focuses only on correctness and often overlooks consistency, necessary to distinguish and address these harms. To bridge this gap, we introduce prompt multiplicity, a framework for quantifying consistency in LLM evaluations. Our analysis reveals significant multiplicity (over 50% inconsistency in benchmarks like Med-HALT), suggesting that hallucination-related harms have been severely misunderstood. Furthermore, we study the role of consistency in hallucination detection and mitigation. We find that: (a) detection techniques detect consistency, not correctness, and (b) mitigation techniques like RAG, while beneficial, can introduce additional inconsistencies. By integrating prompt multiplicity into hallucination evaluation, we provide an improved framework of potential harms and uncover critical limitations in current detection and mitigation strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。