arXiv:2512.02576cs.CV2025-12

用扩散模型生成手势运动,再通过动态图检索合成更自然的同步手势视频。

Co-speech Gesture Video Generation via Motion-Based Graph Retrieval

  • 先用扩散模型学习音频与手势的联合分布,生成合理手势轨迹。
  • 结合高低层音频特征,提升手势生成的上下文准确性。
  • 基于运动相似性检索图路径,拼接不连续节点生成连贯视频。

同步且自然的伴随言语手势视频生成仍是重大挑战。现有方法利用动作图挖掘已有视频数据潜力,但通常依赖输入音频与动作图中动作特征间的距离或共享特征空间嵌入,难以应对音频到手势的多对多映射关系。为此,本文提出新框架:首先使用扩散模型生成手势运动,隐式学习音频与动作的联合分布,从而从输入音频序列生成语境恰当的手势;同时提取音频的低层与高层特征以丰富扩散模型训练。随后,设计精细化的动作基检索算法,通过评估运动的全局与局部相似性,从图中识别最适路径;由于检索路径中的节点未必顺序连续,最后通过无缝拼接各片段生成连贯视频输出。实验表明,该方法在同步精度与手势自然度上显著优于先前方法。

原文摘要 · Abstract (English)

Synthesizing synchronized and natural co-speech gesture videos remains a formidable challenge. Recent approaches have leveraged motion graphs to harness the potential of existing video data. To retrieve an appropriate trajectory from the graph, previous methods either utilize the distance between features extracted from the input audio and those associated with the motions in the graph or embed both the input audio and motion into a shared feature space. However, these techniques may not be optimal due to the many-to-many mapping nature between audio and gestures, which cannot be adequately addressed by one-to-one mapping. To alleviate this limitation, we propose a novel framework that initially employs a diffusion model to generate gesture motions. The diffusion model implicitly learns the joint distribution of audio and motion, enabling the generation of contextually appropriate gestures from input audio sequences. Furthermore, our method extracts both low-level and high-level features from the input audio to enrich the training process of the diffusion model. Subsequently, a meticulously designed motion-based retrieval algorithm is applied to identify the most suitable path within the graph by assessing both global and local similarities in motion. Given that not all nodes in the retrieved path are sequentially continuous, the final step involves seamlessly stitching together these segments to produce a coherent video output. Experimental results substantiate the efficacy of our proposed method, demonstrating a significant improvement over prior approaches in terms of synchronization accuracy and naturalness of generated gestures.

手势生成扩散模型动作图视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。