单步生成真人视频,又快又准,突破实时与质量的矛盾
LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation

- 用布朗桥直接传输图像,避免逐步生成
- 1步生成200帧/秒,长期稳定无失真
- 适合实时视频、虚拟主播等场景
长时长、实时的说话人头像生成仍受延迟-质量权衡限制:多步扩散模型难以流式生成,而实时自回归方法存在误差累积和身份漂移。为此,我们提出LeapTalk,一种新型框架,仅需单步前向即可实现稳定、实时的说话头像生成,并可扩展至任意长度视频。核心是单步桥接蒸馏机制:一方面,摒弃传统噪声到数据范式,基于布朗桥提出数据到数据的传输方式,以持续参考为锚点,有效缓解身份漂移,提升长期时间稳定性;另一方面,通过具有信噪比对齐时间变换Φ(τ)的异构蒸馏框架,实现预训练扩散教师到学生桥模型的平滑知识迁移。此外,提出音频驱动的无分类器引导机制,在极低步数下仍保持精细的唇部同步。大量实验表明,本方法在仅1步条件下达到最高200 FPS,生成视频保真度高、时间一致性强,显著优于现有方法。
原文摘要 · Abstract (English)
Long-form and real-time talking-head generation remains challenging due to a latency-quality trade-off: inefficient multi-step diffusion prohibits streaming generation, whereas real-time autoregressive approaches suffer from error accumulation and identity drift. To address this drawback, we propose LeapTalk, a novel framework that achieves stable and real-time talking-head generation with a single forward step, scaling to arbitrarily long videos. At the heart of our approach lies a single-step bridge distillation scheme. On the one hand, departing from the conventional noise-to-data paradigm, we introduce a data-to-data transport formulation based on a Brownian bridge. Anchored by a persistent reference, this strategy effectively mitigates identity drift and enhances long-term temporal stability. On the other hand, to enable smooth knowledge transfer from a pre-trained diffusion teacher to the student bridge model, we explore a heterogeneous distillation framework with an SNR-aligned time transformation $Φ(τ)$, which bridges the functional discrepancy between the two models. Moreover, we propose an audio-driven classifier-free guidance mechanism to maintain fine-grained lip synchronization under extreme step reduction. Extensive experiments demonstrate that our method achieves high-fidelity and temporally consistent video generation with only 1 step at up to 200 FPS, significantly outperforming existing approaches in both efficiency and stability. Project Page: https://zhangrongxiang.github.io/leaptalk-page/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。