arXiv:2608.03227cs.RO2026-08

用无序姿态数据训练机器人动作流模型,提升复杂动作追踪效果。

PFM-HR: Pose Flow Matching for Humanoid Robots

论文配图:PFM-HR: Pose Flow Matching for Humanoid Robots
图 1 · 摘自论文原文
  • 直接在大规模无序姿态数据上训练流匹配先验,无需顺序动作片段。
  • 引入姿态几何得分,引导策略探索更符合物理结构的动作变化。
  • 适用于高动态动作追踪,可复用于多种任务且不需重新训练先验。

运动先验能提升基于物理的人形机器人追踪的强化学习效果,但时间先验需有序动作片段,而姿态先验对策略诱导的姿态变化指导有限。本文提出面向人形机器人的姿态流匹配(PFM-HR),一种在大规模无序姿态数据上直接训练的可复用流匹配先验。PFM-HR引入姿态几何得分(PGS),量化轨迹中关节坐标变化与先验所捕捉的姿态变化局部几何的一致性。使用PGS调节追踪奖励,引导策略探索趋向结构化姿态变化,同时保持先验在不同追踪任务间冻结不变。实验表明,PFM-HR在单个动作及通用动作追踪中均有提升,尤其在高度动态动作下表现更优。

原文摘要 · Abstract (English)

Motion priors improve reinforcement learning for physics-based humanoid tracking, but temporal priors require ordered motion clips, while pose priors provide limited guidance for policy-induced pose transitions. We present Pose Flow Matching for Humanoid Robots (PFM-HR), a reusable flow matching prior trained directly on large scale unordered pose data. PFM-HR introduces the Pose Geometry Score (PGS), which quantifies how joint coordinate changes during rollouts align with the local geometry of pose variation captured by the prior. Using PGS to modulate the tracking reward guides policy exploration toward structured pose changes while keeping the prior frozen across tracking tasks. Experiments demonstrate that PFM-HR improves both single motion and general motion tracking, especially for highly dynamic motions.

动作生成强化学习姿态追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。