发现大模型藏了大量未被观测的知识,提出新方法量化这些隐藏能力。
Evaluating the Unseen Capabilities: How Many Theorems Do LLMs Know?
- 基于观察到的知识频率外推,估算大模型隐藏知识量。
- 实验显示仅看可见表现会遗漏大量真实知识,低估模型能力。
- 适合评估大模型真实知识储备或比较模型内在能力的研究者使用。
准确评估大语言模型(LLMs)对其能力理解与发展方向至关重要。然而,当前评估常无法真实反映模型实际能力,我们指出其中一个重要原因是忽视了‘未被观测的知识’——即模型内部编码但未在评估中显现的信息。为此,我们提出KnowSum,一种统计框架,通过外推已观察到的知识出现频率,量化特定任务类别的未观测知识量。我们在三个关键应用中验证了KnowSum的有效性:估计总知识量、评估信息检索效果、测量输出多样性。实验表明,仅依赖可观测性能会导致大量知识被遗漏。更重要的是,基于内部知识,KnowSum对多个主流大模型的相对排名产生显著差异。
原文摘要 · Abstract (English)
Accurate evaluation of large language models (LLMs) is crucial for understanding their capabilities and guiding their development. However, current evaluations often inconsistently reflect the actual capacities of these models. In this paper, we demonstrate that one of many contributing factors to this \textit{evaluation crisis} is the oversight of unseen knowledge -- information encoded by LLMs but not directly observed or not yet observed during evaluations. We introduce KnowSum, a statistical framework designed to provide a more comprehensive assessment by quantifying the unseen knowledge for a class of evaluation tasks. KnowSum estimates the unobserved portion by extrapolating from the appearance frequencies of observed knowledge instances. We demonstrate the effectiveness and utility of KnowSum across three critical applications: estimating total knowledge, evaluating information retrieval effectiveness, and measuring output diversity. Our experiments reveal that a substantial volume of knowledge is omitted when relying solely on observed LLM performance. Importantly, KnowSum yields significantly different comparative rankings for several common LLMs based on their internal knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。