arXiv:2608.30165q-bio.QMcs.AI2026-08

用实验循环评估AI在生物领域的科学推理能力

Science sandboxes measure the scientific capability of AI agents

  • 设计可重复的实验-反馈-假设修正循环框架
  • 发现前沿AI能优化指标但缺乏对机制的理解
  • 适合研究AI科学思维与突破认知局限

科学进步不仅依赖于找到解决方案,更在于理解其背后的规律,并据此设计更好的实验。本文提出科学沙盒框架,通过反复的实验、反馈和假设修正,评估AI代理的科学能力。该框架涵盖从真实物理实验到数据驱动模型再到虚构规则的多层次实验方式。我们在此框架下构建了两个生物学场景:基因调控建模与蛋白质适应性预测,并检验了前沿AI模型的表现。结果显示,部分模型虽能优化定量指标,却未能掌握系统内在规则,尤其在规则超出常见生物先验时,科学推理能力明显下降。科学沙盒揭示了这些失败模式,使科学能力边界可测量,并为研究和拓展人工智能的科学探索能力提供了可控环境。

原文摘要 · Abstract (English)

Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from "wet" physical experiments, to "damp" predictive models trained on empirical data, to "dry" invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.

科学智能生物建模推理评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。