arXiv:2505.01746cs.CV2025-05ICLR被引 14

让虚拟人对话时手势自然协调,支持双人实时互动

Co$^{3}$Gesture: Towards Coherent Concurrent Co-speech 3D Gesture Generation with Interactive Diffusion

  • 用双分支扩散模型分别处理两人语音,生成同步手势
  • 新模块捕捉两人动作时间关联,提升交互一致性
  • 首个大规模双人对话手势数据集,适合虚拟人研发

从语音生成手势在虚拟角色动画中已取得显著进展,但现有方法仅关注单人自言自语,忽视双人对话场景下的同步手势建模。为此,我们构建了包含超过700万帧的大型并发共说话手势数据集GES-Inter,涵盖多样化的双人交互姿态序列。同时提出Co³Gesture框架,实现双人协同对话手势的连贯生成。针对两人身体动态不对称的问题,框架基于分离音频条件设计两个协作生成分支。提出时序交互模块(TIM),有效建模双人手势序列间的时序关联作为交互引导,并融合至并发手势生成过程。进一步设计互注意力机制,全面增强交互动作间的依赖学习,从而生成生动连贯的手势。大量实验表明,该方法在新构建的GES-Inter数据集上优于当前最优模型。数据集与代码已公开。

原文摘要 · Abstract (English)

Generating gestures from human speech has gained tremendous progress in animating virtual avatars. While the existing methods enable synthesizing gestures cooperated by individual self-talking, they overlook the practicality of concurrent gesture modeling with two-person interactive conversations. Moreover, the lack of high-quality datasets with concurrent co-speech gestures also limits handling this issue. To fulfill this goal, we first construct a large-scale concurrent co-speech gesture dataset that contains more than 7M frames for diverse two-person interactive posture sequences, dubbed GES-Inter. Additionally, we propose Co$^3$Gesture, a novel framework that enables coherent concurrent co-speech gesture synthesis including two-person interactive movements. Considering the asymmetric body dynamics of two speakers, our framework is built upon two cooperative generation branches conditioned on separated speaker audio. Specifically, to enhance the coordination of human postures with respect to corresponding speaker audios while interacting with the conversational partner, we present a Temporal Interaction Module (TIM). TIM can effectively model the temporal association representation between two speakers' gesture sequences as interaction guidance and fuse it into the concurrent gesture generation. Then, we devise a mutual attention mechanism to further holistically boost learning dependencies of interacted concurrent motions, thereby enabling us to generate vivid and coherent gestures. Extensive experiments demonstrate that our method outperforms the state-of-the-art models on our newly collected GES-Inter dataset. The dataset and source code are publicly available at \href{https://mattie-e.github.io/Co3/}{\textit{https://mattie-e.github.io/Co3/}}.

手势生成双人交互扩散模型虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。