arXiv:2604.21279cs.CV2026-04

用风格码精准编辑人脸属性,不破坏其他特征。

LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation

论文配图:LatRef-Diff: Latent and Reference-Guided Diffusion for Facial Attribute Editing and Style Manipulation
图 1 · 摘自论文原文
  • 用潜空间和参考图生成风格码,替代传统语义方向。
  • 在CelebA-HQ上实现领先效果,支持随机与自定义风格控制。
  • 无需成对数据,通过正反一致性训练提升稳定性。

人脸属性编辑与风格操控在虚拟形象和照片编辑中至关重要。然而,由于面部结构复杂且属性间强相关,精确控制特定属性而不影响其他特征仍具挑战。尽管条件GAN取得进展,但受限于精度与训练不稳。扩散模型虽有潜力,却因语义方向表达力不足而难以实现有效风格操控。本文提出LatRef-Diff,一种基于扩散的新型框架:将传统语义方向替换为风格码,并提出潜空间与参考引导两种生成方式;设计风格调制模块,结合可学习向量、交叉注意力与分层结构,实现高质量图像生成。为提升训练稳定性并消除配对图像需求,引入前后一致性训练策略——先用图像特定语义方向近似移除目标属性,再通过风格调制恢复,由感知损失与分类损失引导。在CelebA-HQ上的大量实验表明,该方法在定性与定量评估中均达到最优性能,消融实验验证了各设计的有效性。

原文摘要 · Abstract (English)

Facial attribute editing and style manipulation are crucial for applications like virtual avatars and photo editing. However, achieving precise control over facial attributes without altering unrelated features is challenging due to the complexity of facial structures and the strong correlations between attributes. While conditional GANs have shown progress, they are limited by accuracy issues and training instability. Diffusion models, though promising, face challenges in style manipulation due to the limited expressiveness of semantic directions. In this paper, we propose LatRef-Diff, a novel diffusion-based framework that addresses these limitations. We replace the traditional semantic directions in diffusion models with style codes and propose two methods for generating them: latent and reference guidance. Based on these style codes, we design a style modulation module that integrates them into the target image, enabling both random and customized style manipulation. This module incorporates learnable vectors, cross-attention mechanisms, and a hierarchical design to improve accuracy and image quality. Additionally, to enhance training stability while eliminating the need for paired images (e.g., before and after editing), we propose a forward-backward consistency training strategy. This strategy first removes the target attribute approximately using image-specific semantic directions and then restores it via style modulation, guided by perceptual and classification losses. Extensive experiments on CelebA-HQ demonstrate that LatRef-Diff achieves state-of-the-art performance in both qualitative and quantitative evaluations. Ablation studies validate the effectiveness of our model's design choices.

人脸编辑扩散模型风格控制无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。