让机器人在复杂接触任务中更可靠地达成目标。
Physics-informed Goal-Conditioned Reinforcement Learning under Hybrid Contact Dynamics

- 引入接触感知与分层结构,有选择地应用物理先验知识
- 在接触密集场景中,传统方法会失效,新方法显著提升性能
- 适合研究机器人操纵、强化学习与物理建模的学者
从稀疏反馈中学习达成任意目标,要求智能体对状态-目标对的可达性有深刻理解。目标条件强化学习(GCRL)通过学习可泛化的策略来应对这一挑战,但当系统动力学变为高维、混合或依赖接触时,泛化能力急剧下降。为此,本文提出物理信息引导的目标条件强化学习(Pi-GCRL),引入最优控制启发的归纳偏置。尽管在导航和无物体任务中表现良好,其在接触密集任务中的可靠性仍不明确——接触交互引发混合动力学、模式依赖可控性及非光滑价值景观。本文分析表明,这些结构特性会导致现有Pi-GCRL方法在接触密集操作中性能退化。基于此,我们提出接触感知与分层形式,有选择地应用物理先验。实验验证了该方法在接触丰富操作任务中的有效性,为扩展Pi-GCRL至复杂操纵提供了原则性路径。
原文摘要 · Abstract (English)
Learning to reach arbitrary goals from sparse feedback requires agents to infer a rich notion of reachability across state--goal pairs. Goal-conditioned reinforcement learning (GCRL) tackles this challenge by learning policies that generalize across goals, but this generalization becomes increasingly difficult as the underlying dynamics become high-dimensional, hybrid, or contact-dependent. To address this issue, physics-informed GCRL (Pi-GCRL) introduces optimal-control-inspired inductive biases into goal-conditioned value learning. While Pi-GCRL methods have proven effective in navigation and object-free goal-reaching domains, their reliability in contact-rich tasks remains unclear, where contact interactions induce hybrid dynamics, mode-dependent controllability, and nonsmooth value landscapes. In this work, we show that these structural properties can cause existing Pi-GCRL methods to degrade when applied naively to contact-rich manipulation. Motivated by this analysis, we introduce contact-aware and hierarchical formulations that apply physics-informed inductive biases selectively across the manipulation problem. Our results provide a principled step toward extending Pi-GCRL to contact-rich manipulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。