arXiv:2505.21701cs.CL2025-05EMNLP被引 9

发现大模型知识探测方法极不稳定,微小改动结果差异巨大。

Do We Know What LLMs Don't Know? A Study of Consistency in Knowledge Probing

  • 用输入扰动+量化指标评估知识探测一致性
  • 同方法内扰动后一致率降至40%,跨方法仅7%一致
  • 适合关注大模型可信度与评测可靠性研究者

大型语言模型(LLMs)的可靠性因幻觉问题严重受损,亟需精准识别其知识盲区。现有知识探测方法包括基于校准和提示两种。本文提出一种新评估流程,通过引入输入变异和量化指标,揭示知识探测存在双重不一致性:(1) 同一方法内部不一致——提示中轻微非语义扰动导致探测结果显著波动,例如仅调换答案选项顺序,一致性即降至约40%;(2) 跨方法不一致——不同探测方法对同一模型、数据集、提示的判断相互矛盾,跨方法决策一致性低至7%。这些发现挑战了现有探测方法的有效性,凸显构建抗扰动探测框架的紧迫性。

原文摘要 · Abstract (English)

The reliability of large language models (LLMs) is greatly compromised by their tendency to hallucinate, underscoring the need for precise identification of knowledge gaps within LLMs. Various methods for probing such gaps exist, ranging from calibration-based to prompting-based methods. To evaluate these probing methods, in this paper, we propose a new process based on using input variations and quantitative metrics. Through this, we expose two dimensions of inconsistency in knowledge gap probing. (1) Intra-method inconsistency: Minimal non-semantic perturbations in prompts lead to considerable variance in detected knowledge gaps within the same probing method; e.g., the simple variation of shuffling answer options can decrease agreement to around 40%. (2) Cross-method inconsistency: Probing methods contradict each other on whether a model knows the answer. Methods are highly inconsistent -- with decision consistency across methods being as low as 7% -- even though the model, dataset, and prompt are all the same. These findings challenge existing probing methods and highlight the urgent need for perturbation-robust probing frameworks.

大模型知识探测一致性幻觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。