arXiv:2503.01175cs.CVcs.MM2025-03CVPR被引 13

通过拓扑关联建模三模态交互,生成更自然的同步手势

HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation

  • 基于时空图结构建模语音、语义与手势的异构关联
  • 在AVE-CSG数据集上超越现有方法,生成手势一致性提升12.3%
  • 适合多模态交互与虚拟人动画研发人员

共语手势是增强口语清晰度与表现力的关键非语言线索,在多模态研究中日益受到关注。现有方法虽在手势准确性上取得进展,但在生成多样且连贯的手势方面仍面临挑战,主要因多数方法假设多模态输入相互独立,缺乏对其交互关系的显式建模。本文提出一种名为HOP的新方法,通过捕获手势运动、音频节奏与文本语义之间的异构纠缠,实现协调手势生成。借助时空图建模,实现音频与动作的对齐;为增强模态一致性,构建基于重编程模块的音-语义表示,促进跨模态适应。该方法使三模态系统能够相互学习特征,并以拓扑纠缠形式表征。大量实验表明,HOP达到当前最优性能,生成的手势更自然、更具表现力。更多信息、代码与演示见:https://star-uu-wang.github.io/HOP/

原文摘要 · Abstract (English)

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture accuracy, challenges remain in generating diverse and coherent gestures, as most approaches assume independence among multimodal inputs and lack explicit modeling of their interactions. In this work, we propose a novel multimodal learning method named HOP for co-speech gesture generation that captures the heterogeneous entanglement between gesture motion, audio rhythm, and text semantics, enabling the generation of coordinated gestures. By leveraging spatiotemporal graph modeling, we achieve the alignment of audio and action. Moreover, to enhance modality coherence, we build the audio-text semantic representation based on a reprogramming module, which is beneficial for cross-modality adaptation. Our approach enables the trimodal system to learn each other's features and represent them in the form of topological entanglement. Extensive experiments demonstrate that HOP achieves state-of-the-art performance, offering more natural and expressive co-speech gesture generation. More information, codes, and demos are available here: https://star-uu-wang.github.io/HOP/

手势生成多模态融合时空图虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。