arXiv:2605.10820cs.AIcs.LG2026-05

评估智能体在物理约束下如何高效测量并发现规律。

MaD Physics: Evaluating information seeking under constraints in physical environments

论文配图:MaD Physics: Evaluating information seeking under constraints in physical environments
图 1 · 摘自论文原文
  • 构建三类物理环境,模拟受预算限制的科学探索
  • 智能体需在有限测量次数内推断隐藏物理定律
  • 适合研究科学推理、约束规划与多模态学习

科学发现本质上是资源受限的过程,需在测量质量与数量间权衡。现有基准或侧重静态知识推理,或忽略实际约束,无法评估真实科学探索能力。为此,我们提出MaD Physics基准,包含基于不同物理定律的三个环境,通过修改物理法则避免先验知识干扰。每个试验中,智能体在耗尽预算前进行测量,随后推断系统背后物理规律以预测未来状态。该基准评估智能体从数据中建模和在约束下规划的核心能力。我们使用四个Gemini模型(2.5 Flash Lite、2.5 Flash、2.5 Pro、3 Flash)进行测试,发现其在结构化探索和数据收集方面存在不足,揭示提升科学推理的关键方向。

原文摘要 · Abstract (English)

Scientific discovery is fundamentally a resource-constrained process that requires navigating complex trade-offs between the quality and quantity of measurements due to physical and cost constraints. Measurements drive the scientific process by revealing novel phenomena to improve our understanding. Existing benchmarks for evaluating agents for scientific discovery focus on either static knowledge-based reasoning or unconstrained experimental design tasks, and do not capture the ability to make measurements and plan under constraints. To bridge this gap, we propose Measuring and Discovering Physics (MaD Physics), a benchmark to evaluate the ability of agents to make informative measurements and conclusions subject to constraints on the quality and quantity of measurements. The benchmark consists of three environments, each based on a distinct physical law. To mitigate contamination from existing knowledge, MaD Physics includes altered physical laws. In each trial, the agent makes measurements of the system until it exhausts an allotted budget and then the agent has to infer the underlying physical law to make predictions about the state of the system in the future. MaD Physics evaluates two fundamental capabilities of scientific agents: inferring models from data and planning under constraints. We also demonstrate how MaD Physics can be used to evaluate other capabilities such as multimodality and in-context learning. We benchmark agents on MaD Physics using four Gemini models (2.5 Flash Lite, 2.5 Flash, 2.5 Pro, and 3 Flash), identifying shortcomings in their structured exploration and data collection capabilities and highlighting directions to improve their scientific reasoning.

科学发现约束规划物理建模智能体评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。