arXiv:2604.17621cs.AI2026-04ACL

测试大模型在结构化知识覆盖与组合推理上的短板

KnowledgeBerg: Evaluating Systematic Knowledge Coverage and Compositional Reasoning in Large Language Models

  • 构建跨10领域17语言的4800道多选题基准
  • 主流模型知识覆盖仅5.26%-36.88%准确,推理准确率16%-44%
  • 揭示模型三类失败模式:缺知识、不知需求、推理错误

许多现实问题看似简单,实则隐含两个能力要求:(i) 对有限知识领域的系统性覆盖,(ii) 在该领域上进行组合式集合推理,我们称之为「冰山一角」现象。通过知识广度(所需知识集基数)与推理深度(组合操作次数)两个正交维度形式化这一挑战。提出KnowledgeBerg基准,包含4,800道多选题,源自1,183个枚举种子,覆盖10个领域和17种语言,知识来源权威可复现。代表性开源LLM在知识枚举任务上仅获5.26-36.88 F1,知识驱动推理任务准确率为16.00-44.19。诊断分析揭示三类失败阶段:完整性缺失、需求意识不足、执行错误。该模式贯穿不同语言与模型规模。尽管测试时计算与检索增强带来显著提升(最高+4.35与+3.78点),仍存在巨大差距,暴露出当前模型在组织结构化知识与执行受限领域组合推理方面的根本局限。数据集已公开于https://huggingface.co/datasets/2npc/KnowledgeBerg。

原文摘要 · Abstract (English)

Many real-world questions appear deceptively simple yet implicitly demand two capabilities: (i) systematic coverage of a bounded knowledge universe and (ii) compositional set-based reasoning over that universe, a phenomenon we term "the tip of the iceberg." We formalize this challenge through two orthogonal dimensions: knowledge width, the cardinality of the required universe, and reasoning depth, the number of compositional set operations. We introduce KnowledgeBerg, a benchmark of 4,800 multiple-choice questions derived from 1,183 enumeration seeds spanning 10 domains and 17 languages, with universes grounded in authoritative sources to ensure reproducibility. Representative open-source LLMs demonstrate severe limitations, achieving only 5.26-36.88 F1 on universe enumeration and 16.00-44.19 accuracy on knowledge-grounded reasoning. Diagnostic analyses reveal three stages of failure: completeness, or missing knowledge; awareness, or failure to identify requirements; and application, or incorrect reasoning execution. This pattern persists across languages and model scales. Although test-time compute and retrieval augmentation yield measurable gains -- up to 4.35 and 3.78 points, respectively -- substantial gaps remain, exposing limitations in how current LLMs organize structured knowledge and execute compositional reasoning over bounded domains. The dataset is available at https://huggingface.co/datasets/2npc/KnowledgeBerg

知识评估组合推理大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。