arXiv:2507.02083cs.AI2025-07NeurIPS被引 8

用虚拟实验平台测试大模型的科学探究能力,发现越复杂系统表现越差。

Measuring Scientific Capabilities of Language Models with a Systems Biology Dry Lab

  • 构建虚拟生物实验环境,模拟真实复杂系统进行迭代实验设计。
  • 在137个小型系统上测试6个前沿大模型,复杂度升高时性能显著下降。
  • 适合研究大模型科学推理能力或系统生物学智能助手的开发者参考。

设计实验与结果解读是生物学中的核心科学能力,研究者通过扰动复杂系统来揭示其内在机制。现有大语言模型(LLM)评估方法因湿实验成本过高(人力、时间、设备)而无法有效测试这些能力。本文提出SciGym,首个面向开放科学发现任务的基准测试,评估LLM在迭代实验设计与分析方面的能力。SciGym通过运行基于系统生物学标记语言(SBML)的虚拟实验环境,实现高效模拟数据生成,成为真实复杂系统的理想测试平台。我们在137个小型系统上评估了6个前沿大模型,并发布了总计350个系统。结果显示,虽更强大的模型表现更优,但所有模型在系统复杂度增加时性能均显著下降,表明当前大模型在科学探究能力方面仍有巨大提升空间。

原文摘要 · Abstract (English)

Designing experiments and result interpretations are core scientific competencies, particularly in biology, where researchers perturb complex systems to uncover the underlying systems. Recent efforts to evaluate the scientific capabilities of large language models (LLMs) fail to test these competencies because wet-lab experimentation is prohibitively expensive: in expertise, time and equipment. We introduce SciGym, a first-in-class benchmark that assesses LLMs' iterative experiment design and analysis abilities in open-ended scientific discovery tasks. SciGym overcomes the challenge of wet-lab costs by running a dry lab of biological systems. These models, encoded in Systems Biology Markup Language, are efficient for generating simulated data, making them ideal testbeds for experimentation on realistically complex systems. We evaluated six frontier LLMs on 137 small systems, and released a total of 350 systems. Our evaluation shows that while more capable models demonstrated superior performance, all models' performance declined significantly as system complexity increased, suggesting substantial room for improvement in the scientific capabilities of LLM agents.

大模型评测系统生物学虚拟实验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。