arXiv:2607.20924cs.CV2026-07

实现高保真妆容迁移,精准控制区域且不改变人脸身份。

MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer

论文配图:MagicMakeup: A Region-Controllable Diffusion Transformer for High-Fidelity Makeup-Transfer
图 1 · 摘自论文原文
  • 通过注意力对齐的区域门控,精确控制妆容编辑范围。
  • 在1024×1024分辨率下实现强区域可控性与高妆容保真度。
  • 适合需要精细妆容编辑的应用,如影视特效与虚拟形象设计。

妆容迁移旨在将参考妆容应用到目标人脸的同时保留原貌身份。尽管基于扩散模型的方法在全脸编辑上取得进展,但区域可控性、妆容保真度和身份一致性仍具挑战。原因包括:(i) 像素与注意力对齐偏差导致干扰扩散至非目标区域,削弱区域控制;(ii) 双图像条件下的转移与保留概念未分离,造成妆容属性与身份耦合;(iii) 缺乏高分辨率、身份一致且带区域标注的数据集以支持细粒度监督。本文提出MagicMakeup,一种基于扩散Transformer的区域可控高保真妆容迁移框架,结合空间约束与概念解耦。为实现精准区域编辑并保留身份,提出像素-注意力对齐的区域门控机制,实施区域特定逻辑门控。为明确转移与保留概念,引入跨模态感知引导,对齐文本与图像特征以增强跨模态概念感知。同时构建1024×1024数据对生成流水线,通过区域特异性妆容移除建立统一合成与真实场景基准。大量定量与定性实验表明,MagicMakeup显著提升区域可控性、妆容保真度与身份一致性,在多种风格、人种与姿态下均表现稳健。

原文摘要 · Abstract (English)

Makeup-transfer applies the reference makeup to the source face while preserving the source identity. Despite advances in full-face editing by diffusion-based methods, strong regional controllability, makeup fidelity, and identity preservation remain challenging. The reasons are (i) pixel-to-attention misalignment that causes spillover into non-target areas and weakens regional control; (ii) unclear transfer/preservation concept separation under two-image conditioning, leading to coupling between makeup attributes and identity; and (iii) the lack of a high-resolution dataset that is identity-consistent and region-labeled for fine-grained supervision. In this paper, we propose MagicMakeup, a diffusion transformer-based framework for region-controllable and high-fidelity makeup transfer, built on spatial constraints and concept disentanglement. To enable precise region-specific editing while preserving identity, we propose Token-Aligned Region Gating, which aligns pixel masks with attention and applies region-specific logit gating. To clarify the concepts of transfer and preservation, we further introduce Cross-Modal Perception Guidance, which aligns text and image features to enhance cross-modal concept perception. We also design a pipeline for the generation of 1024 x 1024 data pairs through region-specific makeup removal and establish a unified benchmark in synthetic and real settings. Extensive quantitative and qualitative experiments show that MagicMakeup improves regional controllability, makeup fidelity, and identity preservation, with strong robustness across styles, races, and poses.

妆容迁移扩散模型区域控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。