arXiv:2505.06436cs.CVcs.AI2025-05

用关键点检测约束生成人脸时的情绪变化,让风格编辑更精准。

My Emotion on your face: The use of Facial Keypoint Detection to preserve Emotions in Latent Space Editing

  • 在损失函数中加入面部关键点检测约束,防止编辑时情绪走样。
  • 实验显示情绪变化减少49%,有效缓解特征纠缠问题。
  • 适合需要固定表情的人脸数据增强,尤其对表情研究有价值。

生成对抗网络如StyleGAN/2可生成逼真人脸图像,并具备语义结构化的潜在空间。现有方法通过在潜在空间中寻找语义方向(如性别、年龄)实现图像编辑,理想情况下仅改变目标特征而保持其他特征不变。然而,特征纠缠导致修改一个属性时不可避免影响表情,限制了其在手势与表情研究中的数据增强应用。为此,本文提出在面部关键点检测模型的损失函数中引入附加约束,通过预训练的关键点检测模型(HFLD)限制表达变化。在原有模型基础上增加该损失项后,定量与定性评估均表明,该方法显著缓解了纠缠问题并维持了面部表情。实验结果显示,情绪变化最多降低49%。与当前最优模型对比,本方法在保留面部姿态与表情的同时实现外观变换,为表情研究提供了可靠的数据增强手段。

原文摘要 · Abstract (English)

Generative Adversarial Network approaches such as StyleGAN/2 provide two key benefits: the ability to generate photo-realistic face images and possessing a semantically structured latent space from which these images are created. Many approaches have emerged for editing images derived from vectors in the latent space of a pre-trained StyleGAN/2 models by identifying semantically meaningful directions (e.g., gender or age) in the latent space. By moving the vector in a specific direction, the ideal result would only change the target feature while preserving all the other features. Providing an ideal data augmentation approach for gesture research as it could be used to generate numerous image variations whilst keeping the facial expressions intact. However, entanglement issues, where changing one feature inevitably affects other features, impacts the ability to preserve facial expressions. To address this, we propose the use of an addition to the loss function of a Facial Keypoint Detection model to restrict changes to the facial expressions. Building on top of an existing model, adding the proposed Human Face Landmark Detection (HFLD) loss, provided by a pre-trained Facial Keypoint Detection model, to the original loss function. We quantitatively and qualitatively evaluate the existing and our extended model, showing the effectiveness of our approach in addressing the entanglement issue and maintaining the facial expression. Our approach achieves up to 49% reduction in the change of emotion in our experiments. Moreover, we show the benefit of our approach by comparing with state-of-the-art models. By increasing the ability to preserve the facial gesture and expression during facial transformation, we present a way to create human face images with fixed expression but different appearances, making it a reliable data augmentation approach for Facial Gesture and Expression research.

人脸生成表情保持关键点检测风格编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。