用视觉理解物理规律,让AI像物理学家一样发现公式。
Mimicking the Physicist's Eye:A VLM-centric Approach for Physics Formula Discovery
- 融合视觉、轨迹与符号推理,模拟科学家发现规律的过程。
- 在5000个实例数据集上,准确率超越现有模型。
- 适合研究物理规律自动发现或多模态推理的学者。
从真实世界观测数据中自动发现物理定律是人工智能的一大挑战。现有方法依赖符号回归或大语言模型,仅处理单模态数据,忽略了物理学家赖以洞察运动现象的丰富视觉表征。这种“感官剥夺”严重削弱了对动态现象中时空模式的识别能力。为此,我们提出VIPER-R1,一种以视觉语言模型为中心的多模态模型,用于基于物理的方程推理。该模型整合视觉感知、轨迹数据与符号推理,模拟科学发现过程。通过课程学习的运动结构诱导(MSI)训练,结合监督微调以解析运动相图,并依据因果思维链(C-CoT)生成假设;随后采用奖励引导的符号校准(RGSC)进行强化学习优化公式结构。推理阶段,训练后的VIPER-R1作为智能体:先提出高置信度的符号假设,再主动调用外部符号回归工具进行符号残差重对齐(SR²),这一过程类比物理学家的微扰分析,使理论模型与实测数据一致。为支持研究,我们构建了包含5000个实例的多模态语料库PhysSymbol。实验表明,VIPER-R1在准确率和可解释性上持续优于当前最先进的视觉语言模型基线,实现了更精确的物理定律发现。
原文摘要 · Abstract (English)
Automated discovery of physical laws from observational data in the real world is a grand challenge in AI. Current methods, relying on symbolic regression or LLMs, are limited to uni-modal data and overlook the rich, visual phenomenological representations of motion that are indispensable to physicists. This "sensory deprivation" severely weakens their ability to interpret the inherent spatio-temporal patterns within dynamic phenomena. To address this gap, we propose VIPER-R1, a multimodal model that performs Visual Induction for Physics-based Equation Reasoning to discover fundamental symbolic formulas. It integrates visual perception, trajectory data, and symbolic reasoning to emulate the scientific discovery process. The model is trained via a curriculum of Motion Structure Induction (MSI), using supervised fine-tuning to interpret kinematic phase portraits and to construct hypotheses guided by a Causal Chain of Thought (C-CoT), followed by Reward-Guided Symbolic Calibration (RGSC) to refine the formula structure with reinforcement learning. During inference, the trained VIPER-R1 acts as an agent: it first posits a high-confidence symbolic ansatz, then proactively invokes an external symbolic regression tool to perform Symbolic Residual Realignment (SR^2). This final step, analogous to a physicist's perturbation analysis, reconciles the theoretical model with empirical data. To support this research, we introduce PhysSymbol, a new 5,000-instance multimodal corpus. Experiments show that VIPER-R1 consistently outperforms state-of-the-art VLM baselines in accuracy and interpretability, enabling more precise discovery of physical laws. Project page: https://jiaaqiliu.github.io/VIPER-R1/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。