arXiv:2505.16520cs.CLcs.AI2025-05ACL被引 11

测试大模型内部状态能否真实编码事实性,发现现有方法难泛化到真实生成数据。

Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs

  • 从表格和问答数据中生成逼真真假句子对,构建更贴近实际的评估数据集。
  • 在两个开源大模型上验证,内部状态对事实性的编码部分有效但泛化能力弱。
  • 为后续事实性研究提供新数据和评估指南,适合模型可信性研究者参考。

事实幻觉是大语言模型(LLMs)面临的主要挑战,会生成不准确或虚构内容,损害可靠性与用户信任。近期研究认为,模型生成错误陈述时,其内部状态会编码真实性信息。然而这些研究多依赖缺乏真实性的合成数据,限制了对模型自生成文本的事实准确性评估的泛化能力。本文通过扩展先前工作,提出:(1)从表格数据中采样合理的真实-虚假事实型句子;(2)基于问答数据集生成依赖于大模型的、真实的真假数据集。对两个开源大模型的分析显示,虽然部分支持以往发现,但将结论推广到模型自生成数据仍具挑战。本研究为未来大模型事实性研究奠定基础,并提供更有效的评估实践指导。

原文摘要 · Abstract (English)

Factual hallucinations are a major challenge for Large Language Models (LLMs). They undermine reliability and user trust by generating inaccurate or fabricated content. Recent studies suggest that when generating false statements, the internal states of LLMs encode information about truthfulness. However, these studies often rely on synthetic datasets that lack realism, which limits generalization when evaluating the factual accuracy of text generated by the model itself. In this paper, we challenge the findings of previous work by investigating truthfulness encoding capabilities, leading to the generation of a more realistic and challenging dataset. Specifically, we extend previous work by introducing: (1) a strategy for sampling plausible true-false factoid sentences from tabular data and (2) a procedure for generating realistic, LLM-dependent true-false datasets from Question Answering collections. Our analysis of two open-source LLMs reveals that while the findings from previous studies are partially validated, generalization to LLM-generated datasets remains challenging. This study lays the groundwork for future research on factuality in LLMs and offers practical guidelines for more effective evaluation.

事实性大模型幻觉检测评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。