arXiv:2608.02642physics.chem-phcs.AI2026-08

评测编码代理在真实分子动力学任务中的表现,发现其可靠性仍有巨大提升空间。

MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows

论文配图:MDArena: Evaluating Coding Agents on Realistic Molecular Dynamics Workflows
图 1 · 摘自论文原文
  • 构建50个来自真实科研项目的容器化分子动力学任务,覆盖29种系统和14类流程。
  • 最佳模型仅48%任务严格通过,膜蛋白系统准备等难题仍未突破。
  • 代理常有部分进展但因细节失误无法复现结果,适合评估自主科研能力。

加速科学发现是人工智能最具影响力的领域之一,而计算生物分子模拟尤为关键。编码代理有望自动化其中大量工作,但其在真实分子动力学(MD)任务中的可靠性尚未充分评估。为此,我们提出MDArena,一个包含50个容器化任务的基准,源自活跃的生物分子模拟项目,涵盖29种分子系统和14类研究协议,包括轨迹分析、复杂系统构建、自由能计算及增强采样。我们评估了六种模型/配置组合,涵盖Codex与OpenCode。其中,Codex GPT-5.5在高推理强度下表现最佳,取得24/50(48%)的严格通过率;中等强度为21/50,OpenCode Gemini Flash 3.5为20/50。所有配置的平均正确率与过程奖励均显著高于严格成功率,表明代理虽有实质性进展,却常因细小错误导致无法复现实验流程。困难任务如膜蛋白系统构建与化学变换自由能设置,对所有配置仍基本无解。因此,MDArena揭示了编码代理作为监督助手与作为自主科研者之间的巨大差距,并提供了一个可复现、可扩展的平台,用于追踪该差距的缩小进程。

原文摘要 · Abstract (English)

Accelerating scientific discovery is among the most consequential applications of AI, and computational biomolecular simulation stands out as a particularly promising target within this broader effort. Coding agents promise to automate significant portions of this workflow, yet their reliability on realistic molecular dynamics (MD) tasks remains poorly characterized. To address this issue, we introduce MDArena, a benchmark of 50 containerized tasks drawn from active biomolecular simulation projects, spanning 29 molecular systems and 14 broad research protocols, including trajectory analysis, complex system preparation, free-energy protocols, and enhanced sampling. We evaluate six model/harness configurations spanning Codex and OpenCode. Among the evaluated configurations, Codex GPT-5.5 at extra-high reasoning effort performs best, reaching 24/50 Strict-Pass@1 successes (48%), followed by Codex GPT-5.5 Medium with 21/50, and OpenCode Gemini Flash 3.5 with 20/50. Average correctness and process rewards are substantially higher than strict success rates across all configurations, indicating that agents frequently make meaningful partial progress but fail on the fine-grained details required for reproducible scientific workflows. Hard tasks remain largely unsolved, particularly membrane-protein system preparation and alchemical free-energy setup, both unsolved or near-unsolved by every evaluated configuration. MDArena thus exposes a substantial gap between the usefulness of coding agents as supervised assistants and their reliability as autonomous MD researchers, while providing a reproducible and extensible platform for tracking progress toward closing it.

分子模拟编码代理基准测试AI科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。