arXiv:2603.04124cs.AIcond-mat.mtrl-sci2026-03被引 4

用可验证奖励训练小模型做结构力学推理,发现它只记模板不真懂物理。

BeamPERL: Parameter-Efficient RL with Verifiable Rewards Specializes Compact LLMs for Structured Beam Mechanics Reasoning

  • 用二值正确奖励和符号求解器,让小模型学物理推理
  • 性能提升66.7%但仅在相似结构有效,换支撑位置就崩溃
  • 中间阶段推理最强,越练越依赖模板,不适合真迁移

我们研究强化学习能否通过精确可验证的奖励,教会小型语言模型进行物理推理,而非仅模式匹配。实验中,使用1.5B参数模型,在梁静力学问题上采用参数高效强化学习(RLVR),以符号求解器提供的二值正确性奖励进行训练,无需教师生成的推理轨迹。最佳检查点在Pass@1指标上相比基础模型提升66.7%。然而,模型能力呈现各向异性:能组合更多载荷,却在支撑位置改变等拓扑变化下失效,尽管需解方程相同。中间检查点推理能力最强,持续优化虽维持高奖励得分,但降低鲁棒性。结果表明,仅靠结果级对齐无法确保可迁移的物理推理;即使奖励信号精确无误,也难以实现方程内化。建议将可验证奖励与结构化推理引导结合,以突破模板匹配局限。

原文摘要 · Abstract (English)

Can reinforcement learning with hard, verifiable rewards teach a compact language model to reason about physics, or does it primarily learn to pattern-match toward correct answers? We study this question by training a 1.5B-parameter reasoning model on beam statics, a classic engineering problem, using parameter-efficient RLVR with binary correctness rewards from symbolic solvers, without teacher-generated reasoning traces. The best BeamPERL checkpoint achieves a 66.7% improvement in Pass@1 over the base model. However, the learned competence is anisotropic: the model generalizes compositionally (more loads) but fails under topological shifts (moved supports) that require the same equilibrium equations. Intermediate checkpoints yield the strongest reasoning, while continued optimization degrades robustness while maintaining reward. These findings reveal a key limitation of outcome-level alignment: reinforcement learning with exact physics rewards induces procedural solution templates rather than internalization of governing equations. The precision of the reward signal - even when analytically exact - does not by itself guarantee transferable physical reasoning. Our results suggest that verifiable rewards may need to be paired with structured reasoning scaffolding to move beyond template matching toward robust scientific reasoning.

强化学习物理推理小模型可验证奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。