arXiv:2608.28802cs.CVcs.AI2026-08

首次大规模评估文本引导模型在人脸编辑中的表现,发现其对发型配饰编辑好但对姿态调整差。

A Large-scale Evaluation of Text-guided Models for Facial Editing

论文配图:A Large-scale Evaluation of Text-guided Models for Facial Editing
图 1 · 摘自论文原文
  • 构建169个面部属性的评测集,系统测试六种文本引导模型在人脸编辑中的表现。
  • 模型在发型和配饰编辑上效果良好,但姿态编辑能力差,普遍存在过度修改现象。
  • 发现模型对深肤色男性和老年人存在显著过度编辑偏见,适合关注公平性研究者参考。

人脸外观编辑支撑了FaceApp和Photoshop等流行应用。生成对抗网络(GAN)和三维可变形模型(3DMM)广泛用于人脸编辑:GAN可实现多样编辑(如改发色、发型),但结果不稳定;3DMM生成稳定编辑,但仅限于姿态与表情调整。近期,如Nano Banana的文本引导扩散模型成为图像编辑的新选择,兼具稳定性与多样性。尽管这些模型已在全场景编辑中被广泛测试(如“让女性弹吉他”),但在人脸编辑领域尚未系统评估。本文首次开展大规模评估(约100万张图像),基于两个主流名人脸数据集(CelebA与CelebSET),对比六种主流文本引导模型在序列化人脸编辑任务中的性能。我们提出Face-Edit-Attributes,目前最大的聚焦发型、配饰与姿态编辑的169项面部属性集合。结果显示,多数模型在发型与配饰编辑上表现良好,但在姿态编辑上表现不佳,且普遍存在过量修改问题(如仅要求改发型却改变发色)。此外,我们评估了各模型的群体偏差,发现几乎所有模型对深肤色男性及年长者均产生更多过量编辑,揭示出显著的公平性风险。代码与数据(含约100万张图像的评测集)已开源。

原文摘要 · Abstract (English)

Facial appearance editing powers popular applications like FaceApp and Photoshop. Generative Adversarial Networks (GANs) and 3D Morphable Models (3DMMs) have been widely used for facial editing. GANs can perform varied facial edits (e.g., changing hair color, hairstyle), but often produce unstable edits. 3DMMs produce stable edits, but can only alter pose and facial expression. Recently, text-guided diffusion models like Nano Banana have become popular for image editing. Text-guided models are a compelling alternative to GANs and 3DMMs since they can produce both stable and varied image edits. While text-guided models have been widely tested for whole-scene edits (e.g., ``make the woman play a guitar''), they have not been comprehensively tested for facial editing. We conducted the first large-scale evaluation ($\sim1$M images evaluated) of six popular text-guided models on a sequential facial editing task. We present Face-Edit-Attributes, the largest collection of $169$ facial editing attributes focused on hair, accessories, and pose edits. We compared model performance using two popular celebrity face datasets: CelebA and CelebSET. Our results show that most models performed hair and accessory edits well, but struggled with editing pose. All models over-edit (e.g., changing hair color when asked only to change the hairstyle). We also evaluated demographic biases in each model. Our results show surprising biases in overediting: almost all models created more overedits for dark-skinned male faces and old faces. The code and data for our results (including our repository of $\sim 1$M images) can be accessed \href{https://github.com/rahul1801/Face-Edit-Bench}{\textcolor{blue}{here}}.

人脸编辑文本引导扩散模型公平性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。