用量化音频生成更顺滑的同步手势,解决抖动和重复问题
SmoothSync: Dual-Stream Diffusion Transformers for Jitter-Robust Beat-Synchronized Gesture Generation from Quantized Audio
- 双流扩散变换器融合音视频特征,提升动作同步性
- 引入抖动抑制损失,使动作平滑度提升62.9%以上
- 通过概率量化实现同一输入生成多样手势,适合虚拟人开发
对话手势生成是合成与语音同步的人类自然手势的关键研究方向。现有方法常存在节奏不一致、动作抖动、脚部滑移及多采样多样性不足等问题。本文提出SmoothSync框架,利用量化音频令牌在新型双流扩散变换器(DiT)架构中生成整体手势并增强采样多样性。具体包括:(1) 通过互补的变压器流融合音视频特征,实现更优同步;(2) 引入抖动抑制损失,提升时间平滑性;(3) 实现概率性音频量化,使相同输入生成不同手势序列。为可靠评估抖动下的节拍同步性,提出抗噪声的平滑节拍一致性指标Smooth-BC。在BEAT2和SHOW数据集上的实验表明,SmoothSync优于当前最优方法:在BEAT2上降低30.6%的FGD、提升10.3% Smooth-BC、增加8.4%多样性,同时抖动减少62.9%,脚部滑移减少17.1%。代码将公开以促进后续研究。
原文摘要 · Abstract (English)
Co-speech gesture generation is a critical area of research aimed at synthesizing speech-synchronized human-like gestures. Existing methods often suffer from issues such as rhythmic inconsistency, motion jitter, foot sliding and limited multi-sampling diversity. In this paper, we present SmoothSync, a novel framework that leverages quantized audio tokens in a novel dual-stream Diffusion Transformer (DiT) architecture to synthesis holistic gestures and enhance sampling variation. Specifically, we (1) fuse audio-motion features via complementary transformer streams to achieve superior synchronization, (2) introduce a jitter-suppression loss to improve temporal smoothness, (3) implement probabilistic audio quantization to generate distinct gesture sequences from identical inputs. To reliably evaluate beat synchronization under jitter, we introduce Smooth-BC, a robust variant of the beat consistency metric less sensitive to motion noise. Comprehensive experiments on the BEAT2 and SHOW datasets demonstrate SmoothSync's superiority, outperforming state-of-the-art methods by -30.6% FGD, 10.3% Smooth-BC, and 8.4% Diversity on BEAT2, while reducing jitter and foot sliding by -62.9% and -17.1% respectively. The code will be released to facilitate future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。