新基准挑战大模型在复杂系统中发现科学定律的真实能力。
NewtonBench: Benchmarking Generalizable Scientific Law Discovery in LLM Agents
- 用反事实定律变化生成324个任务,避免记忆和提升可扩展性。
- 模型发现能力随系统复杂度上升急剧下降,对噪声极敏感。
- 工具辅助反而让强模型过早收敛,适合研究真科学发现的学者。
大语言模型正成为科学定律发现的强大工具,但现有评估基准存在科学相关性、可扩展性和抗记忆性之间的根本权衡。同时,它们将发现简化为静态函数拟合,未能体现通过交互探索复杂系统来揭示隐藏规律的真实科研过程。为此,我们提出NewtonBench,涵盖12个物理领域共324个科学定律发现任务。通过系统性改变经典定律(反事实定律变化),生成大量既科学又难记忆、可扩展的问题。评估从静态拟合升级为交互式模型发现,要求智能体通过实验探测模拟系统以揭示隐藏原理。实验表明,前沿大模型虽具一定发现能力,但随系统复杂度增加迅速退化,且对观测噪声极度敏感。更意外的是,代码解释器等工具可能阻碍强模型——诱导其过早从探索转向利用,满足于次优解。这说明在复杂交互环境中实现鲁棒、泛化的科学发现仍是核心挑战。NewtonBench提供了一个可扩展、稳健且符合科学本质的测试平台,对衡量真实进步和推动下一代具备真正科学发现能力的AI代理至关重要。
原文摘要 · Abstract (English)
Large language models are emerging as powerful tools for scientific law discovery, a foundational challenge in AI-driven science. However, existing benchmarks for this task suffer from a fundamental methodological trilemma, forcing a trade-off between scientific relevance, scalability, and resistance to memorization. Furthermore, they oversimplify discovery as static function fitting, failing to capture the authentic scientific process of uncovering embedded laws through the interactive exploration of complex model systems. To address these critical gaps, we introduce NewtonBench, a benchmark comprising 324 scientific law discovery tasks across 12 physics domains. Our design mitigates the evaluation trilemma by using counterfactual law shifts - systematic alterations of canonical laws - to generate a vast suite of problems that are scalable, scientifically relevant, and memorization-resistant. Moreover, we elevate the evaluation from static function fitting to interactive model discovery, requiring agents to experimentally probe simulated complex systems to uncover hidden principles. Our extensive experiment reveals a clear but fragile capability for discovery in frontier LLMs: this ability degrades precipitously with increasing system complexity and exhibits extreme sensitivity to observational noise. Notably, we uncover a paradoxical effect of tool assistance: providing a code interpreter can hinder more capable models by inducing a premature shift from exploration to exploitation, causing them to satisfice on suboptimal solutions. These results demonstrate that robust, generalizable discovery in complex, interactive environments remains the core challenge. By providing a scalable, robust, and scientifically authentic testbed, NewtonBench offers a crucial tool for measuring true progress and guiding the development of next-generation AI agents capable of genuine scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。