用跨认知域任务评估大模型,更真实反映推理能力。
GIM: Evaluating models via tasks that integrate multiple cognitive domains

- 设计820个原创题目,融合多种认知操作,避免纯记忆或抽象推理。
- 通过53种测试配置、5位评审员,构建稳定的能力评估体系。
- 发现思维预算和量化对性能影响堪比模型选择,且收益递减。
随着大语言模型基准测试趋于饱和,评估界采取两种策略提升难度:一是提高知识要求(如GPQA、HLE),二是去除知识依赖转向抽象推理(如ARC-AGI)。前者混淆记忆与能力,后者使推理脱离实际场景。本文提出一种新方法:基于现实任务的综合评估框架(GIM),包含820个原创问题(615个公开,205个私有),每题需协调约束满足、状态追踪、认知警觉、受众校准等多重认知操作,依托广泛可得的知识,确保推理过程落地真实情境且不依赖专业领域。所有题目均由专家原创,多数配有分项评分标准。我们基于53种测试配置(模型×思考层级组合)、5位校准评审员,利用超过100万次原始评分数据生成的20.38万次平均评分单元,构建了判官感知的连续响应双参数逻辑模型(2PL IRT),获得稳健的能力估计,即使在原始准确率受错误、缺失数据或评审宽松差异干扰时仍能正确排序测试配置。在此框架下,我们呈现涵盖22个模型、47个报告配置的全面排行榜,并开展迄今最广泛的公开研究,揭示在固定基准上测试时计算资源如何权衡模型能力:11个模型覆盖35种配置。结果表明,同一模型家族内的配置选择(如思考预算、量化方式)影响与模型选型相当,且增加思考令牌呈现边际收益递减。
原文摘要 · Abstract (English)
As LLM benchmarks saturate, the evaluation community has pursued two strategies to increase difficulty: escalating knowledge demands (GPQA, HLE) or removing knowledge entirely in favor of abstract reasoning (ARC-AGI). The first conflates memorization with capability; the second divorces reasoning from the practical contexts in which it matters. We take a different approach. The Grounded Integration Measure (GIM) is a benchmark of 820 original problems (615 public, 205 private) where difficulty comes from integration; individual problems require coordinating multiple cognitive operations (constraint satisfaction, state tracking, epistemic vigilance, audience calibration) over broadly accessible knowledge, so that reasoning stays grounded in realistic tasks without being gated on specialized expertise. Each problem is an original expert-authored composition, majority with rubric-decomposed scoring. We calibrate a judge-aware continuous response 2-parameter logistic (2PL) IRT model across 53 test-configurations (unique model x thinking-level pairs) and five calibrated judges, using 203,800 epoch-averaged prompt-judge cells derived from >1M raw judge-scored observations, producing robust ability estimates that correctly order test-configurations even when raw accuracy is distorted by errors, missing data, or judge leniency differences. Using this framework, we present a comprehensive leaderboard spanning 22 models and 47 reporting test-configurations, and conduct what is to our knowledge the most extensive published study of how test-time compute trades off against model capability on a fixed benchmark: 11 models swept across 35 test-configurations. We observe that within-family configuration choices, such as thinking budget and quantization, matter as much as model selection, and increasing thinking tokens has diminishing marginal returns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。