arXiv:2508.16267cs.CLcs.AI2025-08EMNLP被引 4

提出新指标衡量大模型事实性稳定性,发现小模型易崩塌。

From Confidence to Collapse in LLM Factual Robustness

  • 用熵和温度敏感度量化生成过程中的事实稳定性
  • 小模型事实鲁棒分仅0.76,大模型达0.93,不确定性增60%精度下降
  • 适合关注模型可信度与知识持久性的研究者

确保大模型事实知识的鲁棒性对问答与推理等可靠应用至关重要。现有评估多聚焦于性能指标,从提示扰动角度出发,仅反映外部触发的事实鲁棒性。为此,本文提出一种基于生成过程的新方法,通过分析词元分布熵与温度缩放敏感性,构建事实鲁棒性评分(FRS),量化在初始不确定性下事实对解码条件扰动的稳定性。我们在5个大模型上对3个闭卷问答数据集(SQuAD、TriviaQA、HotpotQA)进行广泛实验,结果表明:事实鲁棒性差异显著——小模型的FRS为0.76,大模型达0.93;在不确定性增加时,准确率下降约60%。这些发现揭示了熵与温度缩放对事实准确性的影响,为未来模型的知识保留与检索鲁棒性发展奠定基础。

原文摘要 · Abstract (English)

Ensuring the robustness of factual knowledge in LLMs is critical for reliable applications in tasks such as question answering and reasoning. However, existing evaluation methods predominantly focus on performance-based metrics, often investigating from the perspective of prompt perturbations, which captures only the externally triggered side of knowledge robustness. To bridge this gap, we introduce a principled approach to measure factual robustness from the perspective of the generation process by analyzing token distribution entropy in combination with temperature scaling sensitivity. These two factors build the Factual Robustness Score (FRS), a novel metric which quantifies the stability of a fact against perturbations in decoding conditions, given its initial uncertainty. To validate our approach, we conduct extensive experiments on 5 LLMs across 3 closed-book QA datasets (SQuAD, TriviaQA, and HotpotQA). We show that factual robustness varies significantly -- smaller models report an FRS of $0.76$, larger ones $0.93$ -- with accuracy degrading by ~$60\%$ under increased uncertainty. These insights demonstrate how entropy and temperature scaling impact factual accuracy, and lay a foundation for developing more robust knowledge retention and retrieval in future models.

大模型事实性鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。