arXiv:2412.20251cs.CL2024-12ACL被引 14

通过可控实体频率对比,揭示大模型在低频知识上的脆弱性

ComparisonQA: Evaluating Factuality Robustness of LLMs Through Knowledge Frequency Control and Uncertainty

  • 构建含28.3万题的ComparisonQA基准,成对比较高低频实体表现
  • 发现GPT-4o等模型在低频知识上准确率显著下降,鲁棒性差
  • 利用不确定性识别高质量难题,自动筛选出高难度低频子集

大模型事实知识研究快速发展,现有工作指出其在低频实体相关问题上表现不佳。然而,此类结论不可靠,因问题不仅涉及实体频率差异,还存在难度差异。为此,我们提出ComparisonQA基准,包含28.3万条抽象问题,每题由一个高频与一个低频实体组成,确保仅实体频率不同。同时,结合正确性与不确定性设计双轮评估方法,避免语义捷径问题。实验表明,包括GPT-4o在内的大模型在低频知识上鲁棒性极差。此外,我们发现不确定性可有效识别高质量、无捷径的问题,从而提出自动筛选方法,构建仅含高难度低频问题的ComparisonQA-Hard子集。

原文摘要 · Abstract (English)

The rapid development of LLMs has sparked extensive research into their factual knowledge. Current works find that LLMs fall short on questions around low-frequency entities. However, such proofs are unreliable since the questions can differ not only in entity frequency but also in difficulty themselves. So we introduce ComparisonQA benchmark, containing 283K abstract questions, each instantiated by a pair of high-frequency and low-frequency entities. It ensures a controllable comparison to study the role of knowledge frequency in the performance of LLMs. Because the difference between such a pair is only the entity with different frequencies. In addition, we use both correctness and uncertainty to develop a two-round method to evaluate LLMs' knowledge robustness. It aims to avoid possible semantic shortcuts which is a serious problem of current QA study. Experiments reveal that LLMs, including GPT-4o, exhibit particularly low robustness regarding low-frequency knowledge. Besides, we find that uncertainty can be used to effectively identify high-quality and shortcut-free questions while maintaining the data size. Based on this, we propose an automatic method to select such questions to form a subset called ComparisonQA-Hard, containing only hard low-frequency questions.

大模型评测知识鲁棒性低频知识问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。