用动态组合的知识点构建评测,让大模型难作弊且更考综合理解力。
Encyclo-K: Evaluating LLMs with Dynamically Composed Knowledge Statements
- 以知识语句为单元,测试时随机组合成题,避免记忆数据
- 题目融合8-10个知识点,提升多知识融合评估能力
- 无需专家标注,降低人力成本,适合长期动态评测
评测在追踪大语言模型(LLMs)进展和识别其能力边界中起关键作用。然而现有评测大多在问题层面构建,存在三大根本缺陷:易受数据污染、仅支持单知识点评估、依赖高成本领域专家标注。我们提出Encyclo-K,一种基于语句的评测框架,从底层重构评测设计。核心思想是将知识语句而非问题作为标注单元,问题则由语句动态生成。我们从权威教材中提取独立知识语句,并在测试时通过随机采样动态组合成评测问题。该设计直接解决上述三类问题:组合空间过大难以记忆,模型排名在动态题集间保持稳定,支持定期刷新数据集;每道题整合8-10条语句,实现多知识点综合评估;标注者只需验证格式合规,无需领域知识,大幅降低标注成本。对超过50个LLM的实验表明,Encyclo-K带来显著挑战,具备强区分度。即使顶尖模型OpenAI-GPT-5.1准确率也仅为62.07%,推理模型表现范围为16.04%至62.07%,聊天模型为9.71%至50.40%。结果验证了动态评估与多语句综合理解带来的挑战,确立Encyclo-K作为可扩展的动态评测框架,用于评估模型对细粒度学科知识语句的综合理解能力。
原文摘要 · Abstract (English)
Benchmarks play a crucial role in tracking the rapid advancement of large language models (LLMs) and identifying their capability boundaries. However, existing benchmarks predominantly curate questions at the question level, suffering from three fundamental limitations: vulnerability to data contamination, restriction to single-knowledge-point assessment, and reliance on costly domain expert annotation. We propose Encyclo-K, a statement-based benchmark that rethinks benchmark construction from the ground up. Our key insight is that knowledge statements, not questions, can serve as the unit of curation, and questions can then be constructed from them. We extract standalone knowledge statements from authoritative textbooks and dynamically compose them into evaluation questions through random sampling at test time. This design directly addresses all three limitations: the combinatorial space is too vast to memorize, and model rankings remain stable across dynamically generated question sets, enabling reliable periodic dataset refresh; each question aggregates 8-10 statements for comprehensive multi-knowledge assessment; annotators only verify formatting compliance without requiring domain expertise, substantially reducing annotation costs. Experiments on over 50 LLMs demonstrate that Encyclo-K poses substantial challenges with strong discriminative power. Even the top-performing OpenAI-GPT-5.1 achieves only 62.07% accuracy, and model performance displays a clear gradient distribution--reasoning models span from 16.04% to 62.07%, while chat models range from 9.71% to 50.40%. These results validate the challenges introduced by dynamic evaluation and multi-statement comprehensive understanding. These findings establish Encyclo-K as a scalable framework for dynamic evaluation of LLMs' comprehensive understanding over multiple fine-grained disciplinary knowledge statements.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。