arXiv:2508.08248cs.CV2025-08被引 45

首个可生成无限长语音驱动人像视频的端到端模型

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation

  • 用时步感知音频适配器抑制误差累积
  • 自研音频原生引导机制提升音画同步
  • 适合需要长期稳定语音驱动视频的应用

当前语音驱动人像视频的扩散模型在生成长视频时难以保持自然的音画同步与身份一致性。本文提出 StableAvatar,首个无需后处理即可生成无限长度高质量视频的端到端视频扩散变换器。其基于参考图像和音频输入,通过定制训练与推理模块实现无限长视频生成。我们发现,现有模型无法生成长视频的主要原因是音频建模缺陷:它们依赖第三方现成提取器获取音频嵌入,直接通过交叉注意力注入扩散模型。由于当前扩散主干缺乏音频先验,导致潜空间分布误差随视频片段累积,后续段落潜空间逐渐偏离最优分布。为解决此问题,StableAvatar引入新型时步感知音频适配器,通过时步感知调制防止误差累积。推理时,提出音频原生引导机制,利用扩散模型自身演化中的联合音频-潜空间预测作为动态引导信号,进一步增强音画同步。为提升无限长视频的流畅性,引入动态加权滑动窗口策略,实现潜空间的时序融合。在基准测试上,StableAvatar 在定性和定量评估中均表现出色。

原文摘要 · Abstract (English)

Current diffusion models for audio-driven avatar video generation struggle to synthesize long videos with natural audio synchronization and identity consistency. This paper presents StableAvatar, the first end-to-end video diffusion transformer that synthesizes infinite-length high-quality videos without post-processing. Conditioned on a reference image and audio, StableAvatar integrates tailored training and inference modules to enable infinite-length video generation. We observe that the main reason preventing existing models from generating long videos lies in their audio modeling. They typically rely on third-party off-the-shelf extractors to obtain audio embeddings, which are then directly injected into the diffusion model via cross-attention. Since current diffusion backbones lack any audio-related priors, this approach causes severe latent distribution error accumulation across video clips, leading the latent distribution of subsequent segments to drift away from the optimal distribution gradually. To address this, StableAvatar introduces a novel Time-step-aware Audio Adapter that prevents error accumulation via time-step-aware modulation. During inference, we propose a novel Audio Native Guidance Mechanism to further enhance the audio synchronization by leveraging the diffusion's own evolving joint audio-latent prediction as a dynamic guidance signal. To enhance the smoothness of the infinite-length videos, we introduce a Dynamic Weighted Sliding-window Strategy that fuses latent over time. Experiments on benchmarks show the effectiveness of StableAvatar both qualitatively and quantitatively.

语音驱动视频生成扩散模型无限长视频

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。