让真人视频换脸不失真,文字指令也能精准控制表情动作
DiffMagicFace: Identity Consistent Facial Editing of Real Videos

- 双模型并行控制文本与图像,保持人脸身份一致
- 自建多视角人脸数据集,提升帧间一致性
- 无需视频训练数据,适合复杂动态人脸编辑
文本控制的图像编辑得益于图像扩散模型的发展。然而,将这些技术扩展到真人视频编辑时,如何在视频全程保持面部身份一致、确保编辑主体跨帧连贯成为挑战。本文提出DiffMagicFace,一种融合两个微调后的文本与图像控制模型的视频编辑框架。这两个模型在推理阶段协同工作,生成既保留身份特征又契合编辑语义的视频帧。为保障编辑视频的一致性,我们构建了一个数据集,包含每位被编辑主体的多种面部视角图像,通过渲染技术和优化算法实现。值得注意的是,我们的方法不依赖视频数据集,仍能在一致性与内容质量上达到高质量效果,即使面对说话头像视频和相似类别区分等复杂任务也表现优异。使用该框架生成的视频在视觉效果上可媲美传统渲染软件制作的视频。与当前最先进方法对比,本框架在视觉质量和定量指标上均表现更优。
原文摘要 · Abstract (English)
Text-conditioned image editing has greatly benefitted from the advancements in Image Diffusion Models. However, extending these techniques to facial video editing introduces challenges in preserving facial identity throughout the source video and ensuring consistency of the edited subject across frames. In this paper, we introduce DiffMagicFace, a unique video editing framework that integrates two fine-tuned models for text and image control. These models operate concurrently during inference to produce video frames that maintain identity features while seamlessly aligning with the editing semantics. To ensure the consistency of the edited videos, we develop a dataset comprising images showcasing various facial perspectives for each edited subject. The creation of a data set is achieved through rendering techniques and the subsequent application of optimization algorithms. Remarkably, our approach does not depend on video datasets but still delivers high-quality results in both consistency and content. The excellent effect holds even for complex tasks like talking head videos and distinguishing closely related categories. The videos edited using our framework exhibit parity with videos that are made using traditional rendering software. Through comparative analysis with current state-of-the-art methods, our framework demonstrates superior performance in both visual appeal and quantitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。