arXiv:2512.11423cs.CV2025-12被引 3

实时生成无限时长语音驱动虚拟人,解决延时与质量下降问题

JoyStreamer-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive Diffusion

  • 采用渐进式去噪与运动条件注入,减少误差累积
  • 1.3B参数模型单卡运行达16帧/秒,支持无限长度生成
  • 适合需要低延迟、长时视频生成的交互场景

现有基于DiT的语音驱动虚拟人生成方法虽有进展,但受限于高计算开销和无法生成长视频。自回归方法虽能缓解此问题,却存在误差累积与质量下降。为此,我们提出JoyStreamer-Flash,一种可实时推理且支持无限长度生成的语音驱动自回归模型。其贡献包括:(1)渐进式步数启动(PSB),为初始帧分配更多去噪步骤以稳定生成;(2)运动条件注入(MCI),通过注入带噪声的前帧作为运动条件提升时间一致性;(3)通过缓存重置实现无界旋转位置编码(URCR),支持无限长度生成。该1.3B参数因果模型在单张GPU上实现16 FPS,视觉质量、时间一致性和唇同步表现均具竞争力。

原文摘要 · Abstract (English)

Existing DiT-based audio-driven avatar generation methods have achieved considerable progress, yet their broader application is constrained by limitations such as high computational overhead and the inability to synthesize long-duration videos. Autoregressive methods address this problem by applying block-wise autoregressive diffusion methods. However, these methods suffer from the problem of error accumulation and quality degradation. To address this, we propose JoyStreamer-Flash, an audio-driven autoregressive model capable of real-time inference and infinite-length video generation with the following contributions: (1) Progressive Step Bootstrapping (PSB), which allocates more denoising steps to initial frames to stabilize generation and reduce error accumulation; (2) Motion Condition Injection (MCI), enhancing temporal coherence by injecting noise-corrupted previous frames as motion condition; and (3) Unbounded RoPE via Cache-Resetting (URCR), enabling infinite-length generation through dynamic positional encoding. Our 1.3B-parameter causal model achieves 16 FPS on a single GPU and achieves competitive results in visual quality, temporal consistency, and lip synchronization.

语音驱动自回归生成实时推理无限长度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。