arXiv:2605.04135cs.CYcs.AI2026-05被引 2

AI评估论文常拿旧模型比新前沿,导致能力夸大,误导研究与政策。

Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation

论文配图:Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
图 1 · 摘自论文原文
  • 对比论文测试模型与同期最先进模型的差距,发现平均落后10.85点(ECI)
  • 评估滞后现象每年扩大5.53点,近四分之三源于模型评估延迟
  • 仅3.2%摘要披露推理模式,超一半结论泛化到‘AI’层面,易引发误解

学术界对大模型能力的评估常关注当前可用系统能做什么。但现有文献实际回答的是:几个月或几年前更老、更便宜、更少被深入调用的模型曾做到什么(如2026年论文评测GPT-3.5或GPT-4零样本表现,却对比的是如GPT-5.5 Pro和Claude Opus 4.7这类前沿推理+工具使用系统)。此类研究常缺乏配置细节,结论被抽象为“AI”能力,经引用、媒体和政策传播后形成误判。本研究在预注册下审计112,303条关键词匹配记录(2022年1月至2026年4月),筛选出18,574篇有效论文,其中4,766篇可获取全文。通过复现Arena Elo与Artificial Analysis,基于Epoch AI能力指数(ECI)比较测试模型与同时期前沿水平。结果显示,中位论文评估模型在评估时点落后前沿10.85 ECI(约相当于从Claude Sonnet 3.7到Claude Opus 4.5的距离);探索性理性滞后期模型(H8)分解显示,约25%为同行评审延迟,约75%为评估滞后。该差距以每年+5.53 ECI的速度扩大(95%置信区间[+5.03, +5.83])。此外,仅有3.2%摘要披露推理模式状态,95%置信区间内52.5%的结论将结果泛化至“AI”层面,且该趋势以1.23倍优势/年上升。建议包括提供API补贴、强制编辑要求报告框架(含模型快照、推理模式/投入、工具使用、辅助结构、提示方式等)。推出VERSO-AI 13项清单(核心3项可直接拒稿),扩展现有规范,支持每篇论文的DOI级分析,见frontierlag.org。

原文摘要 · Abstract (English)

Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequentially different, question: what older, cheaper, less-elicited models could do months or years earlier (a 2026 paper evaluating GPT-3.5 or GPT-4 zero-shot, say, against a frontier of reasoning-capable, tool-using systems like GPT-5.5 Pro and Claude Opus 4.7), often reported with sparse configuration details and abstracted upward into claims about "AI" that propagate through citations, media, and policy. We measure the 'publication elicitation gap' (the gap between these answers) in a pre-registered audit of 112,303 LLM-keyword-matched candidate records (2022-01 to 2026-04; 18,574 admissible, 4,766 full-paper texts retrievable), comparing tested models to the contemporaneous frontier on the Epoch AI Capabilities Index (ECI), reproduced under Arena Elo and Artificial Analysis. The median paper evaluates a model +10.85 ECI (~1.4x the distance between Claude Sonnet 3.7 and Claude Opus 4.5) behind the contemporaneous frontier at evaluation time (H1); an exploratory rational-lag baseline (H8) decomposes this into ~25% peer-review latency, ~75% excess lag. The gap is widening at +5.53 ECI/year (H2; 95% CI [+5.03, +5.83]). Meanwhile, only 3.2% of abstracts (21.2% of full-texts) disclose reasoning-mode status on reasoning-capable models (H4) and 52.5% (95% CI [48.2, 56.9]) state conclusions at the level of "AI" rather than the evaluated model(s), rising at OR = 1.23/year. Proposed remedies include API-access subsidies and editorial enforcement of reporting frameworks mandating configuration-surface disclosure (model snapshot, reasoning mode/effort, tool access, scaffolding, prompting, etc.); VERSIO-AI is a 13-item checklist (Core 3 desk-reject) extending existing frameworks at the elicitation surface, with per-DOI analysis at frontierlag.org.

AI评估模型基准能力夸大文献审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。