arXiv:2603.06057cs.CVcs.AI2026-03被引 1

用轻量学生模型实现低延迟语音驱动人脸生成,稳定且对齐更准。

TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head Generation

  • 通过教师-学生蒸馏,少步推理提升生成效率。
  • 在LRS3数据集上实现低延迟,支持边缘设备部署。
  • 引入身份锚定与时间正则化,减少脸型漂移和闪烁。

扩散模型虽已推动逼真人物合成进展,但实际语音驱动人脸生成仍受限于高推理延迟、时间不稳定性(如闪烁、身份漂移)以及复杂语音条件下的音视频不同步。本文提出TempoSyncDiff,一种参考条件的潜在扩散框架,探索少步推理以实现高效语音驱动人脸生成。该方法采用教师-学生蒸馏机制,由标准噪声预测训练的扩散教师指导轻量级学生去噪器,使其能在显著更少的推理步数下运行,提升生成稳定性。框架结合身份锚定与时间正则化,缓解合成过程中的身份漂移与帧间闪烁问题;基于视觉音素(viseme)的音频条件提供粗粒度唇动控制。在LRS3数据集上的实验报告了去噪阶段组件级别的指标(对比VAE重建),并提供了初步延迟分析,包括仅CPU和边缘计算环境下的测量结果及边缘部署可行性评估。结果表明,蒸馏扩散模型可在保留强教师模型大部分重构行为的同时,实现显著更低的延迟推理。本研究定位为在受限计算环境下实现实用化扩散模型人脸生成的初步探索。

原文摘要 · Abstract (English)

Diffusion models have recently advanced photorealistic human synthesis, although practical talking-head generation (THG) remains constrained by high inference latency, temporal instability such as flicker and identity drift, and imperfect audio-visual alignment under challenging speech conditions. This paper introduces TempoSyncDiff, a reference-conditioned latent diffusion framework that explores few-step inference for efficient audio-driven talking-head generation. The approach adopts a teacher-student distillation formulation in which a diffusion teacher trained with a standard noise prediction objective guides a lightweight student denoiser capable of operating with significantly fewer inference steps to improve generation stability. The framework incorporates identity anchoring and temporal regularization designed to mitigate identity drift and frame-to-frame flicker during synthesis, while viseme-based audio conditioning provides coarse lip motion control. Experiments on the LRS3 dataset report denoising-stage component-level metrics relative to VAE reconstructions and preliminary latency characterization, including CPU-only and edge computing measurements and feasibility estimates for edge deployment. The results suggest that distilled diffusion models can retain much of the reconstruction behaviour of a stronger teacher while enabling substantially lower latency inference. The study is positioned as an initial step toward practical diffusion-based talking-head generation under constrained computational settings. GitHub: https://mazumdarsoumya.github.io/TempoSyncDiff

语音驱动低延迟扩散模型人脸生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。