提升人脸分割公平性与鲁棒性,改善生成图像质量
Towards Fair and Robust Face Parsing for Generative AI: A Multi-Objective Approach
- 设计多目标学习框架,动态调整准确率、公平性与鲁棒性的权重
- 在Pix2PixHD生成管道中,使人脸生成更逼真且跨群体一致
- 适用于需要高可靠性的人脸编辑与可控生成场景
人脸分割是计算机视觉中的基础任务,广泛应用于身份验证、人脸编辑和可控图像合成。然而,现有模型常存在公平性不足与鲁棒性差的问题,导致不同人口群体间分割偏差,以及在遮挡、噪声和域偏移下出现错误,进而影响下游生成效果。本文提出一种多目标学习框架,同时优化准确性、公平性与鲁棒性。通过引入基于同伦的损失函数,动态调节各目标权重。在基于Pix2PixHD的GAN生成流水线中对比多目标与单目标U-Net模型,结果表明,具备公平性和鲁棒性的分割显著提升了生成图像的逼真度与一致性。此外,初步实验使用ControlNet(扩散模型的结构化条件模型)探索分割质量对引导生成的影响,验证了多目标分割在提升生成质量方面的有效性。
原文摘要 · Abstract (English)
Face parsing is a fundamental task in computer vision, enabling applications such as identity verification, facial editing, and controllable image synthesis. However, existing face parsing models often lack fairness and robustness, leading to biased segmentation across demographic groups and errors under occlusions, noise, and domain shifts. These limitations affect downstream face synthesis, where segmentation biases can degrade generative model outputs. We propose a multi-objective learning framework that optimizes accuracy, fairness, and robustness in face parsing. Our approach introduces a homotopy-based loss function that dynamically adjusts the importance of these objectives during training. To evaluate its impact, we compare multi-objective and single-objective U-Net models in a GAN-based face synthesis pipeline (Pix2PixHD). Our results show that fairness-aware and robust segmentation improves photorealism and consistency in face generation. Additionally, we conduct preliminary experiments using ControlNet, a structured conditioning model for diffusion-based synthesis, to explore how segmentation quality influences guided image generation. Our findings demonstrate that multi-objective face parsing improves demographic consistency and robustness, leading to higher-quality GAN-based synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。