无需微调,用预训练模型直接生成逼真说话人脸。
IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

- 直接使用Stable Diffusion和IP-Adapter,不需额外微调。
- 唇动同步准确率提升0.16(PCLD),图像质量改善0.7(FID)。
- 适合快速部署、资源有限的研究者使用。
随着扩散模型的快速发展,说话人脸生成取得了显著进展。然而,现有基于扩散模型的方法仍需针对任务进行微调,并依赖大规模音视频数据集,导致计算成本高昂,限制了其在研究社区中的可扩展性和可及性。为此,我们提出一种无需微调的范式,直接利用预训练的Stable Diffusion与IP-Adapter权重实现说话人脸生成。该框架通过IP-Adapter的视觉嵌入能力,从预训练模型中挖掘与口型相关的语义信息。为解决身份漂移、同步误差和时间不稳定性问题,我们设计了三个无训练参数组件:(1) Structurist,显式解耦并重组口型与外观特征,缓解身份漂移和外观失真;(2) Structure Controller,基于准单调运动趋势自适应优化嵌入,实现精准口型同步;(3) Noise Sensor,引入高斯先验检测并抑制闪烁与抖动伪影,增强时间一致性。实验表明,该方法在唇音同步准确率(至少提升0.16,PCLD)和视觉保真度(至少提升0.7,FID)方面均优于现有最先进方法,建立了新的免微调扩散模型说话人脸生成框架。
原文摘要 · Abstract (English)
With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high computational costs that hinder scalability and accessibility of diffusion-based approaches across the research community. To address this, we propose a finetuning-free paradigm that directly performs talking face generation using the pretrained weights of Stable Diffusion and IP-Adapter. This backbone leverages the visual embedding capability of IP-Adapter to mine lip-related semantics from the pretrained Stable Diffusion. To address the challenges of identity drift, synchronization errors, and temporal instability, we also design three trainable-parameterfree components: (1) the Structurist, which explicitly disentangles and reassembles lip and appearance features to mitigate identity drift and appearance distortion; (2) the Structure Controller, which adaptively refines embeddings based on quasi-monotonic motion trends for precise lip synchronization; and (3) the Noise Sensor, which introduces Gaussian prior to detect and suppress flicker and jitter artifacts and enhance temporal consistency. Experimental results show that our method outperforms existing SOTA approaches in both lip-sync accuracy (at least 0.16 gain in PCLD) and visual fidelity (at least 0.7 improvement in FID), establishing a novel fine-tuning-free diffusion framework for talking face generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。