arXiv:2608.03569cs.AIcs.CY2026-08中稿 · ICML

用真实快速变化的领域测试AI科学家的创新能力,发现其核心短板是筛选而非生成。

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities

论文配图:Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities
图 1 · 摘自论文原文
  • 在F1和魔法规则两个真实竞赛领域中评估AI创意能力
  • 最佳模型仅匹配10/40个真实创新(F1),或5/7新卡(MTG)
  • AI擅长生成想法,但缺乏筛选与整合真正有价值的创意能力

评估AI科学家生成新想法的能力极具挑战性。现有基准多依赖合成任务或事后目标,可能受先验知识干扰。本文提出,复杂、对抗性强、快速演进的真实世界领域可作为理想测试场,用于评估推理、新颖性和假设构建等关键能力。研究在两个结构不同的领域中验证:一级方程式赛车(F1)中,模型针对2026赛季赛车设计提出概念,以实际季前创新为真值;魔法规则(MTG)中,模型基于最新卡池设计牌组,并与19个职业巡回赛(PT)牌组对比。结果显示,模型能生成合理方案,但很少与专家实际选择一致。在F1中,最佳模型GPT-5.2在166次生成中仅匹配10个真实创新;在MTG中,最佳牌组来自Gemini 3 Flash,成功复现了第三名PT牌组中的5张新卡,且模型最常选的卡与PT牌组广泛采用的卡高度相关(斯皮尔曼ρ=0.74,p=0.0003)。结果表明,当前AI科学家的核心能力差距在于筛选、优先级排序和连贯的新颖性构建,而非单纯生成。

原文摘要 · Abstract (English)

Benchmarking the ability of AI scientists to generate novel ideas is notoriously difficult. Existing benchmarks in this field have made progress in evaluating scientific reasoning and research replication, but often rely on synthetic tasks or retrospective targets, which may be confounded by prior exposure. We hypothesize that complex, adversarial, fast-moving real-world domains where expert practitioners independently generate observable outputs can provide a practical solution to fill this gap and evaluate the capabilities needed for AI scientists, including reasoning, novelty, and hypothesis formulation. We instantiate this framework in two structurally different domains, Formula 1 (F1), where models ideate around car design concepts for the 2026 season, and real pre-season innovations provide a ground truth, and Magic: The Gathering (MTG), where models propose decks from a recently updated card pool and are evaluated against 19 Pro Tour (PT) decklists. Across both domains, models produce plausible outputs, but few align with real-world expert solutions. In F1, the best model, GPT-5.2 matched 10 of 40 real innovations with 166 ideas proposed across runs. In MTG, the best deck from Gemini 3 Flash recovered 5 of 7 new-set cards from the third-place PT deck, and across all 108 decks, the cards models selected most often were also the cards most widely adopted by PT decks (Spearman $ρ= 0.74$, $p = 0.0003$). These results suggest that a key capability gap for AI scientists is not idea generation, but filtering, prioritization, and coherent novelty.

AI科学家创新评测真实场景筛选能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。