arXiv:2605.02948cs.LGcs.AI2026-05

解决长视频说话头生成中身份漂移问题,实现600秒高保真一致合成。

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation

论文配图:AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation
图 1 · 摘自论文原文
  • 用时序参考编码将静态人脸转为连贯时序表示,缓解时空错位。
  • 采用非对称知识蒸馏,避免自生成参考带来的身份漂移累积。
  • 支持600秒超长视频生成,推理速度达66帧/秒,适合实时应用。

基于扩散模型的说话头生成虽已实现出色视觉质量,但扩展至长时视频仍具挑战。现有分块训练范式存在两大根本缺陷:(1) 静态身份参考与动态音频流间的时空错位;(2) 通过自生成连续参考在各分块间传播的身份漂移。为此,我们提出AsymTalker,一种新型扩散模型方法,包含时序参考编码(TRE)与非对称知识蒸馏(AKD)。首先,TRE通过编码重复时序伪视频,将静态身份图像转化为时序一致的潜在表示,无需额外参数。其次,AKD解决分块训练中的固有条件困境:使用真实参考会导致训练-推理不匹配,而自生成参考则使监督与身份漂移纠缠。我们的非对称设计通过以真实连续参考锚定教师模型,提供无漂移的分块级监督,避免教师瓶颈;同时学生模型仅依赖自生成参考进行推理对齐训练,并通过分布匹配学习,在长时程中保持身份一致性。大量实验表明,AsymTalker在HDTF和VFHQ数据集上达到领先性能,可生成长达600秒、高保真且身份一致的视频,推理速度达66 FPS。

原文摘要 · Abstract (English)

Diffusion-based talking head generation has achieved remarkable visual quality, yet scaling it to long-term videos remains challenging. The widely adopted chunk-wise paradigm introduces two fundamental failures: (1) temporal-spatial misalignment between static identity references and dynamic audio streams, and (2) cascading identity drift propagated through self-generated continuity references across chunks. To address both issues, we propose AsymTalker, a novel diffusion-based talking head generation method comprising Temporal Reference Encoding (TRE) and Asymmetric Knowledge Distillation (AKD). First, TRE mitigates temporal-spatial misalignment by transforming the static identity image into a temporally coherent latent representation through encoding of a temporally replicated pseudo-video, without introducing additional parameters. Second, AKD resolves the inherent conditioning dilemma in chunk-wise training: using ground-truth references causes train-inference mismatch, while self-generated references entangle supervision with identity drift. Our asymmetric design circumvents this by anchoring the teacher model with ground-truth continuity references to provide drift-free, chunk-level supervision, thereby avoiding the teacher bottleneck. Meanwhile, the student model learns under inference-aligned conditions, conditioned only on self-generated references, and is trained via distribution matching to preserve identity over long horizons. Extensive experiments show AsymTalker achieves state-of-the-art results on HDTF and VFHQ. It guarantees high-fidelity, identity-consistent synthesis over 600-second videos and reaches a real-time inference speed of 66 FPS.

说话头生成扩散模型身份一致性长视频合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。