arXiv:2603.14469cs.ROcs.LG2026-03被引 1

让机器人学习更省数据、动作更符合物理规律。

Physics-Informed Policy Optimization via Analytic Dynamics Regularization

  • 用解析力学残差做正则化,直接约束神经策略的输出
  • 在相同任务下,样本效率提升显著,控制更稳定准确
  • 无需改仿真器或强化学习算法,适合现有系统快速集成

强化学习在机器人控制中表现强劲,但当前先进策略学习方法(如演员-评论家)仍存在样本复杂度高、动作物理不一致的问题。这源于神经策略仅从数据中隐式重新发现复杂物理规律,而未利用模拟器中已有的精确动力学模型。本文提出一种名为 PIPER 的新型物理感知强化学习框架,通过解析软物理约束将物理规律无缝嵌入神经策略优化过程。核心在于将可微分的拉格朗日残差作为演员目标函数中的正则化项,该残差来自机器人模拟器描述,能微妙引导策略更新朝向动力学一致解。关键优势是物理约束通过额外损失项实现,无需修改现有模拟器或核心强化学习算法。大量实验表明,该方法显著提升学习效率、稳定性与控制精度,为高效且物理一致的机器人控制树立新范式。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has achieved strong performance in robotic control; however, state-of-the-art policy learning methods, such as actor-critic methods, still suffer from high sample complexity and often produce physically inconsistent actions. This limitation stems from neural policies implicitly rediscovering complex physics from data alone, despite accurate dynamics models being readily available in simulators. In this paper, we introduce a novel physics-informed RL framework, called PIPER, that seamlessly integrates physical constraints directly into neural policy optimization with analytical soft physics constraints. At the core of our method is the integration of a differentiable Lagrangian residual as a regularization term within the actor's objective. This residual, extracted from a robot's simulator description, subtly biases policy updates towards dynamically consistent solutions. Crucially, this physics integration is realized through an additional loss term during policy optimization, requiring no alterations to existing simulators or core RL algorithms. Extensive experiments demonstrate that our method significantly improves learning efficiency, stability, and control accuracy, establishing a new paradigm for efficient and physically consistent robotic control.

强化学习机器人控制物理约束策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。