用扩散模型实现512分辨率高保真唇形同步,解决画质与对齐的矛盾。
HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

- 基于扩散模型端到端生成,直接在512×512分辨率运行
- 在主观质量与同步精度上均达到当前最优水平
- 首次系统消除音频数据泄露问题,让模型真正依赖音频信号
我们提出HighSync,一种基于扩散模型的端到端框架,可生成与任意输入音频高度对齐的逼真说话人脸视频。现有方法普遍难以兼顾图像质量与同步准确性,导致输出或视觉失真,或唇动不同步。HighSync同时解决这两项挑战,据我们所知是首个原生在512×512分辨率运行的唇形同步模型,具备进入影视与广播等专业生产环境的潜力。其核心在于识别并系统性消除此前研究中隐匿的数据泄露现象,该现象阻碍了模型对音频信号的真实依赖。在感知质量与同步精度多项指标上的全面评估表明,HighSync在两项关键性能上均达到当前最优。源代码、预训练模型及补充视频结果已公开:https://github.com/saeed5959/high_sync
原文摘要 · Abstract (English)
We present HighSync, an end-to-end diffusion-based framework for high-fidelity lip synchronization that generates photorealistic talking-face videos aligned with arbitrary input audio. Existing approaches consistently struggle to reconcile image quality with synchronization accuracy, producing either visually degraded outputs or temporally inconsistent lip movements. HighSync addresses both challenges simultaneously and, to our knowledge, is the first lip sync model to operate natively at 512*512 resolution, positioning it as a viable solution for professional production environments such as the film and broadcast industries. Central to our approach is the identification and systematic elimination of a data leakage phenomenon that has silently undermined temporal modeling in prior work, preventing models from developing a genuine dependence on the audio signal. Comprehensive evaluations across both perceptual quality and synchronization accuracy metrics confirm that HighSync achieves state-of-the-art performance on both fronts. Source code, pre-trained models, and supplementary video results are publicly available at: https://github.com/saeed5959/high_sync
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。