arXiv:2412.19009cs.CVcs.MM2024-12被引 14

支持多模态输入的精细人脸局部编辑框架,可保持未编辑区域不变。

FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing

  • 融合草图、文本、颜色图等多模态条件,统一编码到潜在空间
  • 通过自监督隐空间变形纠正对齐误差,实现姿态迁移
  • 支持细粒度语义编辑,适合需要精准控制的图像生成任务

现有面部编辑方法虽取得显著进展,但难以支持多模态条件下的局部编辑。其输出图像质量在多次增量编辑后会显著下降,主要因缺乏局部编辑能力。本文提出一种新型多模态生成与融合框架FACEMUG,实现全局一致的局部面部编辑,支持多种输入模态(如草图、语义图、颜色图、参考图像、文本、属性标签),并能进行细粒度、语义可控的编辑,同时保持未编辑区域不变。通过将所有模态整合至统一的生成潜在空间,设计了新颖的多模态特征融合机制,结合潜在空间和特征空间中的多模态聚合与风格融合模块。进一步引入自监督隐空间变形算法,有效校正面部特征错位,实现编辑图像姿态向给定潜在码的高效迁移。大量实验表明,FACEMUG在编辑质量、灵活性和语义控制方面优于当前最优方法,适用于多种局部人脸编辑任务。

原文摘要 · Abstract (English)

Existing facial editing methods have achieved remarkable results, yet they often fall short in supporting multimodal conditional local facial editing. One of the significant evidences is that their output image quality degrades dramatically after several iterations of incremental editing, as they do not support local editing. In this paper, we present a novel multimodal generative and fusion framework for globally-consistent local facial editing (FACEMUG) that can handle a wide range of input modalities and enable fine-grained and semantic manipulation while remaining unedited parts unchanged. Different modalities, including sketches, semantic maps, color maps, exemplar images, text, and attribute labels, are adept at conveying diverse conditioning details, and their combined synergy can provide more explicit guidance for the editing process. We thus integrate all modalities into a unified generative latent space to enable multimodal local facial edits. Specifically, a novel multimodal feature fusion mechanism is proposed by utilizing multimodal aggregation and style fusion blocks to fuse facial priors and multimodalities in both latent and feature spaces. We further introduce a novel self-supervised latent warping algorithm to rectify misaligned facial features, efficiently transferring the pose of the edited image to the given latent codes. We evaluate our FACEMUG through extensive experiments and comparisons to state-of-the-art (SOTA) methods. The results demonstrate the superiority of FACEMUG in terms of editing quality, flexibility, and semantic control, making it a promising solution for a wide range of local facial editing tasks.

人脸编辑多模态生成局部编辑潜空间对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。