arXiv:2608.22126cs.LGcs.CL2026-08

让AI先建物理模型再算,提升小模型物理推理能力

Decoupled Physical Modeling and Execution for Physics Reasoning

论文配图:Decoupled Physical Modeling and Execution for Physics Reasoning
图 1 · 摘自论文原文
  • 分离物理建模与计算,分两阶段训练强化建模能力
  • 在多个数据集上比GRPO平均提升3%的推理准确率
  • 适合需要精准物理理解的小型视觉语言模型

物理推理需要构建对底层物理系统的统一模型,而非仅依赖符号或公式运算。尽管大语言模型在数学和编程问题上表现优异,但在物理问题上仍受限于建模与计算的纠缠。人类解题时会先建立系统表征再计算。受此启发,我们提出统一框架,通过提炼显式编码物理建模过程的中间表示,并采用两阶段后训练策略:监督微调建立结构化建模,基于评分规则的强化学习提升建模质量。在多个多模态物理基准测试中,该方法显著提升不同模型与数据集上的物理推理性能。在PhysReason、PhyX和SeePhys上,物理建模方法相较GRPO平均提升约3%。结果表明,显式物理建模是提升小型视觉语言模型物理推理能力的有效策略。

原文摘要 · Abstract (English)

Physics reasoning requires constructing a consistent model of the underlying physical system rather than relying solely on symbolic or formula-based manipulation. Although large language models have shown strong ability in solving math and coding problems, they still struggle with physics problems, as these problems entangle the physical modeling process with mathematical calculations. Humans approach physics by first building a representation of the system before performing calculations. Inspired by this, we introduce a unified framework that distills intermediate representations that explicitly encode the physical modeling process and adopt a two-stage post-training strategy, where supervised fine-tuning establishes structured modeling, and reinforcement learning with rubric-based feedback improves the quality of the modeling process. Experiments on multiple multimodal physics benchmarks show that our approach generally improves physical reasoning performance across different models and datasets. Across PhysReason, PhyX, and SeePhys, physical modeling outperforms GRPO by ~3% on average. showing that explicit physical modeling is an effective strategy for improving physics reasoning in small VLMs.

物理推理建模分离小模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。