arXiv:2608.21601cs.AIcs.CL2026-08被引 1

真实科研请求难评估,该研究用实测数据构建新基准,揭示模型表现上限。

K-Bench: measuring model performance on real scientific agent requests

论文配图:K-Bench: measuring model performance on real scientific agent requests
图 1 · 摘自论文原文
  • 从真实用户请求中采样,用九个前沿模型在隔离环境运行1602次任务
  • 仅47.6%的评估达8分门槛,科学准确性平均仅6.22,远低于沟通能力
  • 发现模型普遍夸大成果,强调应关注交付物与宣称的分布关系而非排名

科学人工智能评估多基于可打分任务:选择题、有标准解的代理任务或具已知生成结构的模拟器。但真实科研请求具有未明确定义、附带文件、无真值等特点。我们报告K-Bench 01,一个由K-Dense Web真实用户流量中第一轮请求采样构建的评估体系,由九个前沿模型在相同沙箱中端到端执行,共完成1,602次代理运行。三位盲评语言模型裁判依据八维度评分表对每项结果进行打分。在八维指令明确要求‘领域科学家愿接受微调后成果’的前提下,无一模型在三位裁判中全部达标。gpt-5.6-sol平均分最高(8.04),但其95%置信区间[7.80, 8.23]覆盖阈值,且两位裁判将claude-opus-5列第一。因此我们报告系统排序为可复现量,绝对得分视为评估工具属性,榜首位置保持未定。所有39,934项评分(含八维度分及整体综合评分,不含非适用项)中,47.6%低于8分。评分难度不均:科学准确性均值6.22,通信能力均值7.33,各模型内部方向一致。最常见失败标签为‘过度宣称’,出现在31.4%的评估中。我们认为科学代理的有效信息并非排行榜位置,而是交付内容、宣称内容与产出物之间的联合分布。

原文摘要 · Abstract (English)

Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor instructs judges that a domain scientist would accept the work with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.

科学智能评估基准模型评测真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。