arXiv:2512.11229cs.CVcs.SD2025-12被引 4

提出实时端到端语音驱动人脸生成框架,突破扩散模型速度瓶颈。

REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming Distillation

  • 用压缩潜空间+上下文缓存实现半自回归流式生成
  • 推理速度达40.2帧/秒,生成质量优于现有方法
  • 适合直播、虚拟主播等实时应用

扩散模型显著推动了语音驱动人脸生成(THG)的发展。然而,推理速度慢和普遍采用的非自回归范式严重制约了基于扩散模型的THG应用。本文提出REST,首个基于扩散模型的实时端到端流式语音驱动人脸生成框架。为支持实时端到端生成,首先通过高压缩比的时空变分自编码器学习紧凑的视频潜空间。为在紧凑潜空间内实现半自回归流式生成,引入ID-上下文缓存机制,将身份-锚点与上下文缓存原理融合于键值缓存,以保持长期生成中的身份一致性和时序连贯性。此外,提出异步流式蒸馏(ASD)策略,利用具有异步噪声调度的非流式教师模型监督流式学生模型,缓解误差累积并提升时序一致性。REST弥合了自回归与扩散方法间的差距,在实时THG应用中实现效率突破。实验表明,REST在生成速度和整体性能上均优于当前最优方法。

原文摘要 · Abstract (English)

Diffusion models have significantly advanced the field of talking head generation (THG). However, slow inference speeds and prevalent non-autoregressive paradigms severely constrain the application of diffusion-based THG models. In this study, we propose REST, a pioneering diffusion-based, real-time, end-to-end streaming audio-driven talking head generation framework. To support real-time end-to-end generation, a compact video latent space is first learned through a spatiotemporal variational autoencoder with a high compression ratio. Additionally, to enable semi-autoregressive streaming within the compact video latent space, we introduce an ID-Context Cache mechanism, which integrates ID-Sink and Context-Cache principles into key-value caching for maintaining identity consistency and temporal coherence during long-term streaming generation. Furthermore, an Asynchronous Streaming Distillation (ASD) strategy is proposed to mitigate error accumulation and enhance temporal consistency in streaming generation, leveraging a non-streaming teacher with an asynchronous noise schedule to supervise the streaming student. REST bridges the gap between autoregressive and diffusion-based approaches, achieving a breakthrough in efficiency for applications requiring real-time THG. Experimental results demonstrate that REST outperforms state-of-the-art methods in both generation speed and overall performance.

人脸生成扩散模型实时生成流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。