提出KScope框架,精准判断大模型对问题的真实掌握程度。
KScope: A Framework for Characterizing the Knowledge Status of Language Models
- 构建五类知识状态分类体系,通过分层统计检验逐级推断模型知识模式。
- 发现外部上下文能缩小模型间知识差距,且相关特征显著提升知识更新效果。
- 揭示模型在部分错误与冲突时行为相似,但持续错误时差异巨大,适合评估模型可靠性。
刻画大语言模型对特定问题的知识状态极具挑战性。以往研究多聚焦于模型内部参数记忆与外部上下文冲突下的表现,却未能充分反映模型真实掌握程度。本文首先提出基于一致性与正确性的五类知识状态分类体系;进而提出KScope——一种分层统计检验框架,可逐步细化关于知识模式的假设,并将模型知识归类至五种状态之一。我们在九个大模型上对四个数据集进行系统评估,发现:(1)支持性上下文可缩小模型间知识差距;(2)难度、相关性与熟悉度等上下文特征驱动有效知识更新;(3)当模型部分正确或存在冲突时行为相似,但持续错误时差异显著;(4)基于特征分析约束的上下文摘要结合增强可信度,进一步提升更新效果并实现跨模型泛化。
原文摘要 · Abstract (English)
Characterizing a large language model's (LLM's) knowledge of a given question is challenging. As a result, prior work has primarily examined LLM behavior under knowledge conflicts, where the model's internal parametric memory contradicts information in the external context. However, this does not fully reflect how well the model knows the answer to the question. In this paper, we first introduce a taxonomy of five knowledge statuses based on the consistency and correctness of LLM knowledge modes. We then propose KScope, a hierarchical framework of statistical tests that progressively refines hypotheses about knowledge modes and characterizes LLM knowledge into one of these five statuses. We apply KScope to nine LLMs across four datasets and systematically establish: (1) Supporting context narrows knowledge gaps across models. (2) Context features related to difficulty, relevance, and familiarity drive successful knowledge updates. (3) LLMs exhibit similar feature preferences when partially correct or conflicted, but diverge sharply when consistently wrong. (4) Context summarization constrained by our feature analysis, together with enhanced credibility, further improves update effectiveness and generalizes across LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。