通过动态交互建模,提升3D多人运动预测的结构一致性和社交合理性。
Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

- 基于残差流匹配与确定性粗略先验,稳定生成骨架序列。
- 动态跨主体交互机制随时间同步消息传递,提升多人协作真实性。
- 适合关注多人动作生成、社交行为建模的研究者使用。
3D多人运动预测需同时建模个体运动规律与人与人之间的交互关系。尽管流匹配能有效生成多假设以提高预测精度,但直接从纯噪声预测骨骼序列会损害结构一致性,并在早期噪声主导阶段引入不可靠的跨主体交互。为此,本文提出一种先验引导的残差流匹配框架:首先,采用确定性粗略先验(DCP)建立运动学锚点,将生成过程形式化为对运动残差的条件流,简化生成目标并保持结构稳定性;其次,设计动态跨交互(DCI)机制,使主体间消息传递与生成进度在时间上同步,确保可靠社会上下文提取,提升多人运动保真度;最后,采用解耦关节运动架构与双向融合策略,有效保留细粒度运动连贯性。大量实验表明,该方法在多个数据集上达到当前最优性能。代码已开源。
原文摘要 · Abstract (English)
3D multi-person motion prediction requires modeling both individual kinematics and inter-person interactions. While Flow Matching is effective for multi-hypothesis generation to improve prediction accuracy, directly predicting skeletal sequences from pure noise often compromises structural consistency and introduces unreliable cross-agent interactions during early noise-dominated integration steps. To address this, we propose a Prior-Guided Residual Flow Matching framework. First, a Deterministic Coarse Prior (DCP) establishes a kinematic anchor, formulating the generative process as a conditional flow over motion residuals to simplify the generative objective and preserve structural stability. Second, a Dynamic Cross-Interaction (DCI) mechanism temporally synchronizes inter-agent message-passing with the integration progress, ensuring the extraction of reliable social contexts and improving multi-person motion fidelity. Finally, a decoupled joint-motion architecture with bidirectional fusion effectively preserves fine-grained kinematic coherence. Extensive experiments demonstrate that our approach achieves state-of-the-art prediction accuracy across multiple datasets. Code is available at https://github.com/Wei-Wei-a/Residual-Flow-Matching-with-Dynamic-Cross-Interaction-for-3D-Multi-Person-Motion-Prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。