提升语音驱动头像视频的对齐与质量,解决细节失真和音画不一致问题。
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
- 分区域动态加权训练损失,增强关键部位生成质量。
- 引入音频-视觉分类器,显著改善语音与动作一致性。
- 支持高效推理与高质量输出,适合虚拟主播等应用。
语音驱动人类动画广泛应用于人机交互,扩散模型的出现进一步推动了该技术发展。现有方法多依赖多阶段生成与中间表示,导致推理时间长,且在特定前景区域生成质量差、音画一致性不足,根源在于缺乏局部细粒度监督。为此,本文提出分区域感知语音驱动上身动画框架(PAHA),包含两个核心方法:基于姿态置信度动态调整区域损失权重的分区域重加权(PAR),以及构建并训练基于扩散模型的区域级音视频分类器以增强运动与共说话音一致性的分区域一致性增强(PCE)。随后设计两种新型推理引导策略:顺序引导(SG)与差分引导(DG),分别平衡效率与质量。此外,构建首个公开的中文新闻主播语音数据集CNAS,以推动该领域研究与验证。大量实验与用户研究证明,PAHA在音画对齐与视频评估指标上显著优于现有方法。代码与数据集将在录用后发布。
原文摘要 · Abstract (English)
Audio-driven human animation technology is widely used in human-computer interaction, and the emergence of diffusion models has further advanced its development. Currently, most methods rely on multi-stage generation and intermediate representations, resulting in long inference time and issues with generation quality in specific foreground regions and audio-motion consistency. These shortcomings are primarily due to the lack of localized fine-grained supervised guidance. To address above challenges, we propose Parts-aware Audio-driven Human Animation, PAHA, a unit enhancement and guidance framework for audio-driven upper-body animation. We introduce two key methods: Parts-Aware Re-weighting (PAR) and Parts Consistency Enhancement (PCE). PAR dynamically adjusts regional training loss weights based on pose confidence scores, effectively improving visual quality. PCE constructs and trains diffusion-based regional audio-visual classifiers to improve the consistency of motion and co-speech audio. Afterwards, we design two novel inference guidance methods for the foregoing classifiers, Sequential Guidance (SG) and Differential Guidance (DG), to balance efficiency and quality respectively. Additionally, we build CNAS, the first public Chinese News Anchor Speech dataset, to advance research and validation in this field. Extensive experimental results and user studies demonstrate that PAHA significantly outperforms existing methods in audio-motion alignment and video-related evaluations. The codes and CNAS dataset will be released upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。