用强化学习提升头戴设备动作恢复的细节精度
MotionGRPO: Overcoming Low Intra-Group Diversity in GRPO-Based Egocentric Motion Recovery

- 通过分组相对策略优化增强扩散模型采样多样性
- 在多个数据集上达到最佳视觉保真度,局部关节误差显著降低
- 适合做精细动作重建的研究者和开发者参考
本文研究从头戴设备信号中恢复全身3D人体运动。现有基于扩散模型的方法依赖全局分布匹配,导致局部关节重建误差。我们提出MotionGRPO,一种利用强化学习后训练注入细粒度引导的新框架。技术上,将扩散采样建模为马尔可夫决策过程,通过分组相对策略优化(GRPO)进行求解。为此,引入混合奖励机制,结合学习的条件感知模型以保证全局视觉合理性,以及显式约束以提升局部关节精度。关键洞察是:扩散恢复中的策略优化因组内样本多样性不足而面临梯度消失问题。为此,进一步提出噪声注入策略,显式增加样本方差,稳定学习过程。大量实验表明,MotionGRPO在多个基准上实现领先性能,具备更优的视觉保真度。
原文摘要 · Abstract (English)
This paper studies full-body 3D human motion recovery from head-mounted device signals. Existing diffusion-based methods often rely on global distribution matching, leading to local joint reconstruction errors. We propose MotionGRPO, a novel framework leveraging reinforcement learning post-training to inject fine-grained guidance into the diffusion process. Technically, we model diffusion sampling as a Markov decision process optimized via Group Relative Policy Optimization (GRPO). To this end, we introduce a hybrid reward mechanism that combines a learned conditioned perceptual model for global visual plausibility and explicit constraints for local joint precision. Our key technical insight is that policy optimization in diffusion-based recovery suffers from vanishing gradients due to limited intra-group sample diversity. To address this, we further introduce a noise-injection strategy that explicitly increases sample variance and stabilizes learning. Extensive experiments demonstrate that MotionGRPO achieves state-of-the-art performance with superior visual fidelity
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。