提出主动感知策略,让模型更智能地学习人体运动规律。
Breaking the Passive Learning Trap: An Active Perception Strategy for Human Motion Prediction
- 用商空间表示法解耦运动几何与坐标冗余
- 通过掩码和噪声注入实现主动学习,提升预测精度
- 可适配多种模型,适合做动作预测的研究者
3D人体运动预测是人工代理对人类行为进行细粒度理解与认知的重要体现。现有方法过度依赖隐式建模时空关系与运动特征,陷入被动学习陷阱,导致获取冗余单调的3D坐标信息,缺乏主动引导的显式学习机制。为此,我们提出主动感知策略(APS),利用商空间表征显式编码运动属性,并引入辅助学习目标强化时空建模。首先设计数据感知模块,将姿态投影至商空间,解耦运动几何与坐标冗余;通过联合编码切向量与格拉斯曼投影,实现几何降维、语义解耦与动态约束强化,有效表征运动姿态。进一步设计网络感知模块,通过恢复性学习主动学习时空依赖:故意遮蔽特定关节或注入噪声,构建辅助监督信号;专用辅助学习网络主动适应并学习扰动信息。显著的是,APS具有模型无关性,可集成于不同预测模型以增强主动感知能力。实验表明,该方法达到新最优性能,在H3.6M上提升16.3%,在CMU Mocap上提升13.9%,在3DPW上提升10.1%。
原文摘要 · Abstract (English)
Forecasting 3D human motion is an important embodiment of fine-grained understanding and cognition of human behavior by artificial agents. Current approaches excessively rely on implicit network modeling of spatiotemporal relationships and motion characteristics, falling into the passive learning trap that results in redundant and monotonous 3D coordinate information acquisition while lacking actively guided explicit learning mechanisms. To overcome these issues, we propose an Active Perceptual Strategy (APS) for human motion prediction, leveraging quotient space representations to explicitly encode motion properties while introducing auxiliary learning objectives to strengthen spatio-temporal modeling. Specifically, we first design a data perception module that projects poses into the quotient space, decoupling motion geometry from coordinate redundancy. By jointly encoding tangent vectors and Grassmann projections, this module simultaneously achieves geometric dimension reduction, semantic decoupling, and dynamic constraint enforcement for effective motion pose characterization. Furthermore, we introduce a network perception module that actively learns spatio-temporal dependencies through restorative learning. This module deliberately masks specific joints or injects noise to construct auxiliary supervision signals. A dedicated auxiliary learning network is designed to actively adapt and learn from perturbed information. Notably, APS is model agnostic and can be integrated with different prediction models to enhance active perceptual. The experimental results demonstrate that our method achieves the new state-of-the-art, outperforming existing methods by large margins: 16.3% on H3.6M, 13.9% on CMU Mocap, and 10.1% on 3DPW.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。