arXiv:2502.10729cs.CV2025-02被引 1

通过视觉风格信息提升语音伴随手势生成的多样性与自然度。

VarGes: Improving Variation in Co-Speech 3D Gesture Generation via StyleCLIPS

  • 引入风格参考视频提取风格特征,增强输入信息。
  • 在多个数据集上显著提升手势多样性和动作自然度。
  • 适合虚拟人、动画和人机交互场景使用。

从音频生成富有表现力且多样的人类手势在人机交互、虚拟现实和动画领域至关重要。现有方法虽取得显著进展,但受限于数据集多样性不足及音频输入信息有限。为此,我们提出 VarGes,一种基于风格线索的变异性驱动框架,以提升语音伴随3D手势生成质量。首先,通过变异性增强特征提取模块(VEFE)将风格参考视频融入3D人体姿态估计网络,提取风格片段(StyleCLIPS),丰富输入风格信息。随后,采用变异性补偿风格编码器(VCSE),基于变压器结构与加性注意力池化层,稳健编码多样化风格表示并有效处理风格差异。最后,变异性驱动手势预测模块(VDGP)通过交叉注意力融合MFCC音频特征与风格编码,注入跨条件自回归模型,实现基于音频与风格线索的3D手势生成。在基准数据集上的实验表明,该方法在手势多样性和自然度方面均优于现有方法。代码与视频结果将在接受后公开:https://github.com/mookerr/VarGES/。

原文摘要 · Abstract (English)

Generating expressive and diverse human gestures from audio is crucial in fields like human-computer interaction, virtual reality, and animation. Though existing methods have achieved remarkable performance, they often exhibit limitations due to constrained dataset diversity and the restricted amount of information derived from audio inputs. To address these challenges, we present VarGes, a novel variation-driven framework designed to enhance co-speech gesture generation by integrating visual stylistic cues while maintaining naturalness. Our approach begins with the Variation-Enhanced Feature Extraction (VEFE) module, which seamlessly incorporates \textcolor{blue}{style-reference} video data into a 3D human pose estimation network to extract StyleCLIPS, thereby enriching the input with stylistic information. Subsequently, we employ the Variation-Compensation Style Encoder (VCSE), a transformer-style encoder equipped with an additive attention mechanism pooling layer, to robustly encode diverse StyleCLIPS representations and effectively manage stylistic variations. Finally, the Variation-Driven Gesture Predictor (VDGP) module fuses MFCC audio features with StyleCLIPS encodings via cross-attention, injecting this fused data into a cross-conditional autoregressive model to modulate 3D human gesture generation based on audio input and stylistic clues. The efficacy of our approach is validated on benchmark datasets, where it outperforms existing methods in terms of gesture diversity and naturalness. The code and video results will be made publicly available upon acceptance:https://github.com/mookerr/VarGES/ .

手势生成风格迁移3D姿态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。