arXiv:2503.21847cs.GRcs.AI2025-03

让语音同步的肢体动作更自然真实,支持零样本生成。

ReCoM: Realistic Co-Speech Motion Generation with Recurrent Embedded Transformer

  • 用递归嵌入变压器建模语音与动作的时空关联。
  • 动作真实度提升86.7%,FGD从18.70降到2.48。
  • 适合做虚拟人、动画生成,无需额外标注数据。

我们提出ReCoM,一个高效框架,用于生成与语音同步的高保真、可泛化的肢体动作。核心创新是递归嵌入变压器(RET),在视觉变换器(ViT)架构中引入动态嵌入正则化(DER),显式建模语音-动作动态关系。该结构实现联合时空依赖建模,通过连贯运动合成提升手势自然度与保真度。为增强鲁棒性,引入DER策略,使模型具备抗噪和跨域泛化能力,显著提升未见语音输入下的零样本生成自然度与流畅性。为缓解自回归推理中的误差累积与自我修正能力弱的问题,提出迭代重构推理(IRI)策略:通过循环姿态重建,结合两类关键组件——无分类器引导提升生成与真实手势分布对齐,无需辅助监督;时序平滑过程消除帧间突变,保障运动连续性。在基准数据集上的大量实验验证了有效性,各项指标达当前最优。特别地,弗雷歇手势距离(FGD)从18.70降至2.48,真实度提升86.7%。

原文摘要 · Abstract (English)

We present ReCoM, an efficient framework for generating high-fidelity and generalizable human body motions synchronized with speech. The core innovation lies in the Recurrent Embedded Transformer (RET), which integrates Dynamic Embedding Regularization (DER) into a Vision Transformer (ViT) core architecture to explicitly model co-speech motion dynamics. This architecture enables joint spatial-temporal dependency modeling, thereby enhancing gesture naturalness and fidelity through coherent motion synthesis. To enhance model robustness, we incorporate the proposed DER strategy, which equips the model with dual capabilities of noise resistance and cross-domain generalization, thereby improving the naturalness and fluency of zero-shot motion generation for unseen speech inputs. To mitigate inherent limitations of autoregressive inference, including error accumulation and limited self-correction, we propose an iterative reconstruction inference (IRI) strategy. IRI refines motion sequences via cyclic pose reconstruction, driven by two key components: (1) classifier-free guidance improves distribution alignment between generated and real gestures without auxiliary supervision, and (2) a temporal smoothing process eliminates abrupt inter-frame transitions while ensuring kinematic continuity. Extensive experiments on benchmark datasets validate ReCoM's effectiveness, achieving state-of-the-art performance across metrics. Notably, it reduces the Fréchet Gesture Distance (FGD) from 18.70 to 2.48, demonstrating an 86.7% improvement in motion realism. Our project page is https://yong-xie-xy.github.io/ReCoM/.

动作生成语音同步扩散模型自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。