arXiv:2507.15550cs.LGcs.AI2025-07NeurIPS被引 8

测试大模型在物理探索中如何利用先验知识,区分不同复杂度下的表现。

PhysGym: Benchmarking LLMs in Interactive Physics Discovery with Controlled Priors

  • 通过可控先验知识设计交互式物理环境,评估大模型推理能力。
  • 不同先验条件下,模型在复杂任务中的假设准确率差异显著。
  • 适合研究科学发现、具身智能与大模型认知机制的学者使用。

评估基于大语言模型的智能体在科学发现方面的能力,特别是其应对环境复杂性变化及利用先验知识的能力,目前缺乏专门的基准。为此,我们提出 extsc{PhysGym},一个新型基准套件与仿真平台,用于严格评估大模型在交互式物理环境中的科学推理能力。其核心贡献在于对提供给智能体的先验知识水平进行精细控制,使研究人员能够沿着问题复杂度和先验知识水平等维度拆解智能体的表现。该基准包含一系列交互式模拟任务,智能体需主动探测环境,在约束条件下逐步收集数据,并推断潜在物理规律。 extsc{PhysGym} 提供标准化的评估协议与指标,用于衡量假设准确性与模型保真度。我们通过基线大模型的结果展示了该基准的实用性,验证了其在不同先验和任务复杂度下区分模型能力的有效性。

原文摘要 · Abstract (English)

Evaluating the scientific discovery capabilities of large language model based agents, particularly how they cope with varying environmental complexity and utilize prior knowledge, requires specialized benchmarks currently lacking in the landscape. To address this gap, we introduce \textsc{PhysGym}, a novel benchmark suite and simulation platform for rigorously assessing LLM-based scientific reasoning in interactive physics environments. \textsc{PhysGym}'s primary contribution lies in its sophisticated control over the level of prior knowledge provided to the agent. This allows researchers to dissect agent performance along axes including the complexity of the problem and the prior knowledge levels. The benchmark comprises a suite of interactive simulations, where agents must actively probe environments, gather data sequentially under constraints and formulate hypotheses about underlying physical laws. \textsc{PhysGym} provides standardized evaluation protocols and metrics for assessing hypothesis accuracy and model fidelity. We demonstrate the benchmark's utility by presenting results from baseline LLMs, showcasing its ability to differentiate capabilities based on varying priors and task complexity.

科学发现大模型评测交互式环境先验知识

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。