arXiv:2503.23888cs.CVcs.AI2025-03被引 3

用文本生成精细掩码,实现高可控的面部编辑

MuseFace: Text-driven Face Editing via Diffusion-based Mask Generation Approach

  • 通过文本驱动的扩散模型生成语义掩码
  • 实现高保真、精细化的面部编辑效果
  • 适合需要精准控制的个性化图像处理场景

面部编辑在个性化图像定制与增强中起关键作用。尽管已有大量工作在文本驱动的面部编辑上取得显著进展,但现有方法仍难以同时满足多样性、可控性与灵活性。为此,我们提出MuseFace,一种完全依赖文本提示的文本驱动面部编辑框架。该框架结合文本到掩码的扩散模型与语义感知的面部编辑模型,可直接从文本生成细粒度语义掩码并完成编辑。文本到掩码扩散模型提供多样性和灵活性,语义感知编辑模型保障可控性。该框架能生成精细语义掩码,实现精准面部编辑,显著提升模型的可控性与灵活性。大量实验表明,MuseFace在高保真度方面表现优异。

原文摘要 · Abstract (English)

Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in text-driven face editing, they still face significant challenges as none of them simultaneously fulfill the characteristics of diversity, controllability and flexibility. To address this challenge, we propose MuseFace, a text-driven face editing framework, which relies solely on text prompt to enable face editing. Specifically, MuseFace integrates a Text-to-Mask diffusion model and a semantic-aware face editing model, capable of directly generating fine-grained semantic masks from text and performing face editing. The Text-to-Mask diffusion model provides \textit{diversity} and \textit{flexibility} to the framework, while the semantic-aware face editing model ensures \textit{controllability} of the framework. Our framework can create fine-grained semantic masks, making precise face editing possible, and significantly enhancing the controllability and flexibility of face editing models. Extensive experiments demonstrate that MuseFace achieves superior high-fidelity performance.

面部编辑扩散模型文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。