用统一框架生成双人交互动作,支持文本、音乐等多种输入
Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow
- 基于修正流实现确定性快速采样,比扩散模型更高效
- 引入检索增强生成,提升动作与条件语义的匹配度
- 适合虚拟伴侣、机器人和游戏角色的动作生成
生成真实且情境感知的双人动作仍面临挑战。现实应用如虚拟/增强现实伙伴、社交机器人和游戏角色,要求模型能生成协调的互动行为,并灵活切换交互与反应模式。我们提出DualFlow,首个统一高效的多模态双人动作生成框架。该框架支持文本、音乐及先前动作序列等多种输入。利用修正流技术,实现从噪声到数据的确定性直线采样路径,显著降低推理时间并减少扩散模型中常见的误差累积。为增强语义对齐,引入新颖的检索增强生成(RAG)模块:通过音乐特征和大语言模型解析空间关系、身体动作与节奏模式,检索动作范例。采用对比式修正流目标强化条件信号对齐,加入同步损失提升两人间时间协调性。在交互、反应及多模态基准上广泛评估表明,DualFlow持续提升动作质量、响应速度与语义一致性,达到当前最优性能,生成连贯、富有表现力且节奏同步的动作。
原文摘要 · Abstract (English)
Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and game agents require models capable of producing coordinated interpersonal behaviour while flexibly switching between interactive and reactive generation. We introduce DualFlow, the first unified and efficient framework for multi-modal two-person motion generation. DualFlow conditions 3D motion generation on diverse inputs, including text, music, and prior motion sequences. Leveraging rectified flow, it achieves deterministic straight-line sampling paths between noise and data, reducing inference time and mitigating error accumulation common in diffusion-based models. To enhance semantic grounding, DualFlow employs a novel Retrieval-Augmented Generation (RAG) module for two-person motion that retrieves motion exemplars using music features and LLM-based text decompositions of spatial relations, body movements, and rhythmic patterns. We use a contrastive rectified flow objective to further sharpen alignment with conditioning signals and add synchronisation loss to improve inter-person temporal coordination. Extensive evaluations across interactive, reactive, and multi-modal benchmarks demonstrate that DualFlow consistently improves motion quality, responsiveness, and semantic fidelity. DualFlow achieves state-of-the-art performance in two-person multi-modal motion generation, producing coherent, expressive, and rhythmically synchronized motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。