arXiv:2512.21373cs.SEcs.AI2025-12被引 6

测试大模型在真实科研代码库中完成完整科研开发任务的能力

AInsteinBench: Benchmarking Coding Agents on Scientific Repositories

  • 从6个主流科研代码库的维护者提交的合并请求中提取真实开发任务
  • 通过多阶段筛选和专家评审确保任务科学性与难度适配
  • 关注模型在可执行环境中的真实开发能力,非单纯生成代码

我们提出AInsteinBench,一个大规模基准测试,用于评估大语言模型(LLM)代理是否能在真实的科研软件生态中担任科学计算开发角色。与侧重概念知识的科学推理基准或强调通用功能实现的软件工程基准不同,AInsteinBench在生产级科研代码库的真实开发场景下评估模型表现。该基准包含来自量子化学、量子计算、分子动力学、数值相对论、流体动力学和化学生物信息学等六个广泛使用的科学代码库的任务,所有任务均基于维护者提交的拉取请求,并经过多阶段筛选与专家评审,确保任务具备科学挑战性、充分的测试覆盖和合理难度。通过在可执行环境中评估、识别科学意义的失败模式以及测试驱动验证,AInsteinBench衡量模型是否能超越表面代码生成,具备计算科学研究所需的核心能力。

原文摘要 · Abstract (English)

We introduce AInsteinBench, a large-scale benchmark for evaluating whether large language model (LLM) agents can operate as scientific computing development agents within real research software ecosystems. Unlike existing scientific reasoning benchmarks which focus on conceptual knowledge, or software engineering benchmarks that emphasize generic feature implementation and issue resolving, AInsteinBench evaluates models in end-to-end scientific development settings grounded in production-grade scientific repositories. The benchmark consists of tasks derived from maintainer-authored pull requests across six widely used scientific codebases, spanning quantum chemistry, quantum computing, molecular dynamics, numerical relativity, fluid dynamics, and cheminformatics. All benchmark tasks are carefully curated through multi-stage filtering and expert review to ensure scientific challenge, adequate test coverage, and well-calibrated difficulty. By leveraging evaluation in executable environments, scientifically meaningful failure modes, and test-driven verification, AInsteinBench measures a model's ability to move beyond surface-level code generation toward the core competencies required for computational scientific research.

代码生成科研自动化大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。