用大模型检测文本幻觉,精准定位错误信息。
Fine-Grained Detection of Context-Grounded Hallucinations Using LLMs
- 提出基于自由文本描述的新幻觉表示方法,覆盖所有可能错误类型。
- 构建超千例人工标注基准,最优模型F1仅0.67,验证任务难度。
- 揭示大模型误判主因:错标缺失细节为矛盾,难识别源外正确信息。
上下文无关的幻觉指模型输出中无法在源文本中验证的信息。本文研究大模型在局部化此类幻觉上的适用性,作为现有复杂评估流程的更实用替代方案。由于缺乏针对幻觉定位的元评估基准,我们构建了一个专为大模型设计的基准,包含超过1,000个挑战性的人工标注样本。同时开发了基于大模型的评估协议,并通过人工评估验证其有效性。鉴于现有幻觉表征方式限制了可表达错误类型,我们提出一种基于自由文本描述的新表征方法,能够完整捕捉各类错误。我们对四个大规模大模型进行了全面评估,结果显示该基准极具挑战性,最优模型的F1分数仅为0.67。通过深入分析,我们提炼出该任务的最佳提示策略,并识别出大模型面临的主要困难:(1) 尽管被指示仅检查输出中的事实,仍倾向于将缺失细节错误标记为不一致;(2) 难以处理含有源文本未提及但事实正确的信息——这类信息虽真实,却无法验证,仅与模型参数知识对齐。
原文摘要 · Abstract (English)
Context-grounded hallucinations are cases where model outputs contain information not verifiable against the source text. We study the applicability of LLMs for localizing such hallucinations, as a more practical alternative to existing complex evaluation pipelines. In the absence of established benchmarks for meta-evaluation of hallucinations localization, we construct one tailored to LLMs, involving a challenging human annotation of over 1,000 examples. We complement the benchmark with an LLM-based evaluation protocol, verifying its quality in a human evaluation. Since existing representations of hallucinations limit the types of errors that can be expressed, we propose a new representation based on free-form textual descriptions, capturing the full range of possible errors. We conduct a comprehensive study, evaluating four large-scale LLMs, which highlights the benchmark's difficulty, as the best model achieves an F1 score of only 0.67. Through careful analysis, we offer insights into optimal prompting strategies for the task and identify the main factors that make it challenging for LLMs: (1) a tendency to incorrectly flag missing details as inconsistent, despite being instructed to check only facts in the output; and (2) difficulty with outputs containing factually correct information absent from the source - and thus not verifiable - due to alignment with the model's parametric knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。