构建首个覆盖261个细分学科的权威评测基准,解决大模型知识评估三大痛点。
Knowledge Index of Noah's Ark

- 以专家锚点为依据,用贪心近似法实现学科代表性量化
- 顶模得分53.17%,但整体仍有显著提升空间,呈分层结构
- 引入激励相容机制与稳定性统计,避免小样本排名误读
大语言模型知识评测面临三大问题:规模驱动设计难以体现学科代表性;简单评分易导致敷衍共识;在有限测试预算下排名不稳定。本文提出KINA,一个涵盖261个细粒度学科、共899项题目的基准。首先,将代表性建模为专家标注锚点的覆盖目标,并通过代理指标实现(1-1/e)贪心近似(命题1),该保证适用于代理指标而非总体代表性。其次,证明奖金奖励制锦标赛在发布评审质量上弱于平价支付的FOSD主导性,且激励相容阈值B > ΔC / Δp_min(定理1)。对13家机构的42个模型评估显示,顶级模型Gemini-3.1-Pro-Preview得分为53.17%,次之Claude-Opus-4.6为49.92%,GPT-5.4为48.55%,远未达饱和。完整榜单呈现分层结构:前沿梯队超48%,主力模型集中在38%-45%区间,低分模型仅略高于10%随机基线。工具增强在五项测评中带来最高5.17分提升,增益因模型而异。报告了自助法排名稳定性统计,显式揭示有限预算下的方差,防止对相邻排名过度解读。
原文摘要 · Abstract (English)
Knowledge benchmarks for LLMs face three issues: scaling-driven designs that do not operationalize disciplinary representativeness; flat-payment annotation that permits lazy consensus; and unaudited ranking instability under bounded test budgets. We introduce KINA, an 899-item benchmark across 261 fine-grained disciplines, with two formal results. First, we cast representativeness as a coverage-style objective over expert-elicited anchors and operationalize disciplinary representativeness through a proxy, yielding a (1-1/e) greedy approximation (Proposition 1); the guarantee applies to the proxy, not to population representativeness. Second, we prove a bonus-on-bar tournament weakly FOSD-dominates flat payment in released-review quality, with incentive-compatibility threshold B > Delta C / Delta p_min (Theorem 1). Evaluating 42 models from 13 labs, the top model, Gemini-3.1-Pro-Preview, reaches 53.17%, followed by Claude-Opus-4.6 at 49.92% and GPT-5.4 at 48.55%, leaving substantial headroom below saturation. The full leaderboard shows a tiered structure rather than a smooth total order: a small frontier tier lies above 48%, a dense strong-model tier spans roughly 38-45%, and low-performing models remain only modestly above the 10% chance baseline. Tool augmentation adds up to 5.17 points across the five tool-use evaluations, with gains varying substantially across models. We report bootstrap ranking-stability statistics to make bounded-budget variance explicit and to discourage over-interpretation of adjacent ranks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。