通过分段解码与门控融合,提升人体动作预测的时序建模能力。
Enhancing Human Motion Prediction via Multi-range Decoupling Decoding with Gating-adjusting Aggregation
- 分阶段解码:按不同未来时长分离特征学习,捕捉多尺度运动模式。
- 动态融合:根据输入数据自适应调整各阶段贡献,提升预测准确性。
- 通用性强:可无缝接入现有方法,显著改善预测性能。
姿态序列的表达对准确建模人体动作预测至关重要。尽管基于深度学习的方法在学习运动表示方面表现良好,但这些方法往往忽视历史信息与未来时刻之间相关性的变化——短期预测相关性更强,远期则较弱,这限制了运动表示的学习并影响预测效果。本文提出一种新方法:多范围解耦解码与门控调整聚合(MD2GA),利用时间相关性优化运动表示学习。该方法采用两阶段策略:第一阶段通过多范围解耦解码,将共享特征分解为不同未来长度的输出,不同解码器提供多样化的运动模式洞察;第二阶段通过门控调整聚合,根据输入动作数据动态融合这些洞察。大量实验表明,该方法可轻松集成到其他预测模型中,并有效提升预测性能。
原文摘要 · Abstract (English)
Expressive representation of pose sequences is crucial for accurate motion modeling in human motion prediction (HMP). While recent deep learning-based methods have shown promise in learning motion representations, these methods tend to overlook the varying relevance and dependencies between historical information and future moments, with a stronger correlation for short-term predictions and weaker for distant future predictions. This limits the learning of motion representation and then hampers prediction performance. In this paper, we propose a novel approach called multi-range decoupling decoding with gating-adjusting aggregation ($MD2GA$), which leverages the temporal correlations to refine motion representation learning. This approach employs a two-stage strategy for HMP. In the first stage, a multi-range decoupling decoding adeptly adjusts feature learning by decoding the shared features into distinct future lengths, where different decoders offer diverse insights into motion patterns. In the second stage, a gating-adjusting aggregation dynamically combines the diverse insights guided by input motion data. Extensive experiments demonstrate that the proposed method can be easily integrated into other motion prediction methods and enhance their prediction performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。