arXiv:2501.18898cs.CVcs.GR2025-01ICCV被引 45

用流匹配加速语音同步手势生成,提升自然度与效率。

GestureLSM: Latent Shortcut based Co-Speech Gesture Generation with Spatial-Temporal Modeling

  • 通过时空注意力显式建模身体各部位间交互关系。
  • 采用流匹配机制,推理速度比扩散模型快10倍以上。
  • 适合数字人、虚拟助手等实时交互场景应用。

基于语音生成全身动作仍面临质量和速度的挑战。现有方法将身体不同部位(如躯干、四肢、手部)分别建模,无法捕捉其空间关联,导致动作不自然且断裂。同时,自回归或扩散模型需数十步推理,生成速度慢。为此,我们提出GestureLSM,一种基于流匹配的语音同步手势生成方法,结合时空建模。该方法通过空间和时间注意力显式建模分块身体区域间的交互,生成连贯的全身动作;引入流匹配机制,显式建模潜在速度空间,实现高效采样。为克服流匹配基线性能不足,我们提出隐空间捷径学习与β分布时间戳采样策略,提升生成质量并加速推理。在BEAT2数据集上达到当前最优性能,推理速度显著优于现有方法,展现出在数字人与具身智能体中部署的潜力。

原文摘要 · Abstract (English)

Generating full-body human gestures based on speech signals remains challenges on quality and speed. Existing approaches model different body regions such as body, legs and hands separately, which fail to capture the spatial interactions between them and result in unnatural and disjointed movements. Additionally, their autoregressive/diffusion-based pipelines show slow generation speed due to dozens of inference steps. To address these two challenges, we propose GestureLSM, a flow-matching-based approach for Co-Speech Gesture Generation with spatial-temporal modeling. Our method i) explicitly model the interaction of tokenized body regions through spatial and temporal attention, for generating coherent full-body gestures. ii) introduce the flow matching to enable more efficient sampling by explicitly modeling the latent velocity space. To overcome the suboptimal performance of flow matching baseline, we propose latent shortcut learning and beta distribution time stamp sampling during training to enhance gesture synthesis quality and accelerate inference. Combining the spatial-temporal modeling and improved flow matching-based framework, GestureLSM achieves state-of-the-art performance on BEAT2 while significantly reducing inference time compared to existing methods, highlighting its potential for enhancing digital humans and embodied agents in real-world applications. Project Page: https://andypinxinliu.github.io/GestureLSM

手势生成流匹配数字人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。