让大模型在虚构物理世界中自主发现规律,检验真实科学推理能力。
DiscoverPhysics: Benchmarking LLMs for Out-of-the-Box Scientific Thinking

- 设计22个违背现实物理的虚拟世界,逼大模型通过实验推导新定律。
- 顶尖模型仅能正确破解一半世界,尤其在隐藏结构探测上表现不佳。
- 适合评估大模型科学探究能力,尤其关注实验设计与假设修正能力。
前沿大语言模型在多项物理测评中表现优异,但难以区分真正的科学推理与已有知识的回忆。我们提出DiscoverPhysics,一个交互式基准测试,要求大模型代理在模拟世界中发现其运动规律,而这些世界故意偏离真实物理。构建了22个世界,涵盖屏蔽引力、分数幂引力、多粒子耦合、类暗物质粒子、非坐标无关物理及随时间变化的相互作用等。每个世界由N体模拟器按需生成,代理需设计多轮实验,观察原始轨迹数据,最终提交自然语言解释和推断定律的Python实现。由于解题需设计有信息量的实验并不断修正假设,该基准考察长时程推理能力。评估基于两个互补维度:未见粒子的轨迹均方误差(MSE)和基于专家制定评分标准的LLM判分解释质量。在11个前沿模型中,最强模型仅通过一半世界,且在需揭示隐含结构的世界中持续失败。开源模型在实验设计和结论提取上显著落后于商业模型。进一步发现,预测精度高并不保证解释质量好,概念理解依赖于通过精心设计的实验进行假设迭代优化。
原文摘要 · Abstract (English)
Frontier LLMs now perform strongly across a wide range of physics evaluations, but it is hard to disentangle genuine reasoning from recall of established science. We introduce DiscoverPhysics, an interactive benchmark that asks a LLM agent to discover the laws of motion of a simulated world whose physics deliberately deviates from our own. We construct 22 worlds governed by, among others, screened and fractional-power gravity, multi-species couplings, hidden dark-matter-like particles, non-coordinate-free physics, and time-varying interactions. Each world is generated on demand by an N-body simulator, for which the agent proposes several rounds of experiments, observes raw trajectory data, and ultimately submits both a natural-language explanation of the world's physics and a Python implementation of the inferred law. Because solving a world requires the agent to design informative experiments and revise its hypotheses, the benchmark probes long-horizon reasoning over an experimental history. We evaluate submissions along two complementary axes: trajectory MSE on held-out particles and an LLM-judged explanation score following an expert-written rubric assessing conceptual understanding of each world. Across eleven frontier models, we find that the strongest agents pass only half of the worlds and consistently fail on those where latent structure must be uncovered. Open-source models lag substantially behind commercial models, both in their ability to design informative experiments and in extracting conclusions from the data. We further find that good predictive accuracy does not guarantee high explanation quality and that conceptual understanding depends on hypothesis refinement through well-chosen experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。