用自动研究框架评测模型如何自主改进世界模型。
AutoWorldModel-Bench: A State-Centric Benchmark for Automated World-Model Research

- 设定固定算力下,让代码代理自主优化基础世界模型。
- 64次实验中59次提升,超半数显著增益(Δ≥+0.10)。
- 适合评估前沿代码代理在开放科研任务中的表现。
世界建模仍是一个未定型领域:架构、训练目标与状态表示之间存在复杂交互,尚无通用方案适用于所有环境。这使其成为人工智能编码代理作为自主研究人员的理想测试场——不同于当前主流基准中预先指定目标的工程任务,此处改进方向不预先设定。我们提出 AutoWorldModel-Bench,一个闭环基准,前沿代码代理在固定算力预算下自主改进给定的基础世界模型。该基准涵盖八种游戏环境,采用统一的结构化状态表示——从每款游戏提取的真实实体状态,通过共享张量格式输入,从而将动态建模与感知分离,实现每轮运行仅需几分钟。在64次实验中,Codex-5.4和Claude Opus 4.6在保留测试集上全部取得正向提升(除一次外),其中33次提升幅度达+0.10及以上,其余亦为正值;91%的胜出修改涉及模型或训练机制的根本性调整,而非超参数微调。该基准提供了一个可评估前沿代码代理在开放式科研任务中能力的场景。
原文摘要 · Abstract (English)
World modeling is an unsettled field: architectures, training objectives, and state representations interact in complex ways, and no single recipe dominates across environments. This makes it an ideal testbed for AI coding agents acting as autonomous researchers--a setting in which the improvement direction is not specified in advance, unlike the engineering-to-spec tasks that dominate current agent benchmarks. We introduce AutoWorldModel-Bench, a closed-loop benchmark in which frontier coding agents autonomously improve a provided base world model under a fixed compute budget. The benchmark spans eight game environments under a unified structured-state representation--ground-truth entity state extracted from each game and consumed through a shared tensor format--which isolates dynamics modeling from perception and enables minutes-per-run iteration. Across 64 sessions, Codex-5.4 and Claude Opus 4.6 improve their base on a held-out test split in all but one session, with about half (33 of 64) a substantial gain ($Δ\geq +0.10$) and the remaining improvements smaller but positive; in 91% of sessions the winning edit is a substantive change to the model or training rather than a hyperparameter tweak. Our benchmark offers a setting in which frontier coding agents can be evaluated on open-ended research rather than engineering-to-spec problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。