用扩散模型实现保脸型的精准人脸视频编辑
IP-FaceDiff: Identity-Preserving Facial Video Editing with Diffusion
- 基于预训练扩散模型微调,实现文本驱动局部编辑
- 编辑速度提升80%,视频帧间保持连贯性
- 支持多样表情与姿态变化,适合内容创作者使用
人脸视频编辑对内容创作日益重要,可操控面部表情与特征。现有模型存在编辑质量差、计算成本高、难以跨多种编辑保持脸型一致等问题,且常受限于预定义属性,灵活性不足。为此,我们提出一种新框架,利用预训练文本到图像扩散模型的丰富隐空间,针对性微调用于人脸视频编辑任务。该方法引入定向微调策略,在实现高质量、局部化、文本驱动编辑的同时,确保视频帧间身份一致性。通过在推理中复用预训练模型,编辑时间显著减少80%,并保持视频序列的时序连贯性。我们在多种挑战场景下进行评估,涵盖不同头姿、复杂动作序列和多样表情。结果表明,该方法在多项指标与基准测试中持续优于现有技术。
原文摘要 · Abstract (English)
Facial video editing has become increasingly important for content creators, enabling the manipulation of facial expressions and attributes. However, existing models encounter challenges such as poor editing quality, high computational costs and difficulties in preserving facial identity across diverse edits. Additionally, these models are often constrained to editing predefined facial attributes, limiting their flexibility to diverse editing prompts. To address these challenges, we propose a novel facial video editing framework that leverages the rich latent space of pre-trained text-to-image (T2I) diffusion models and fine-tune them specifically for facial video editing tasks. Our approach introduces a targeted fine-tuning scheme that enables high quality, localized, text-driven edits while ensuring identity preservation across video frames. Additionally, by using pre-trained T2I models during inference, our approach significantly reduces editing time by 80%, while maintaining temporal consistency throughout the video sequence. We evaluate the effectiveness of our approach through extensive testing across a wide range of challenging scenarios, including varying head poses, complex action sequences, and diverse facial expressions. Our method consistently outperforms existing techniques, demonstrating superior performance across a broad set of metrics and benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。