arXiv:2509.12176cs.LG2025-09

用对抗学习实现无需配对数据的高保真人脸操作。

From Autoencoders to CycleGAN: Robust Unpaired Face Manipulation via Adversarial Learning

  • 基于循环GAN,融合身份与感知损失,保持人脸特征一致。
  • 在无配对数据下,真实度、结构保留和身份相似性均优于自编码器。
  • 适合需要跨姿态光照变化的人脸生成场景,如影视特效。

人脸合成与操控在娱乐和人工智能中日益重要,尤其在仅能获取无配对、无对齐数据的情况下,仍需生成高度真实且身份不变的图像。本文研究基于对抗学习的无配对人脸操控,从自编码器基线出发,构建鲁棒的指导式循环生成对抗网络(CycleGAN)框架。自编码器虽能捕捉粗略身份特征,但常丢失细节。本方法引入谱归一化提升训练稳定性,采用身份感知与感知损失以维持主体身份和高层结构,并结合关键点加权的循环约束,确保姿态与光照变化下的面部几何一致性。实验表明,该对抗训练的循环生成模型在真实度(FID)、感知质量(LPIPS)和身份保留度(ID-Sim)上均优于自编码器,循环重建相似性(SSIM)表现良好,推理速度实用,在无配对数据下达到高质量效果,接近在精心整理的配对子集上使用pix2pix的效果。结果证明,指导式谱归一化循环生成网络为从自编码器迈向稳健的无配对人脸操控提供了可行路径。

原文摘要 · Abstract (English)

Human face synthesis and manipulation are increasingly important in entertainment and AI, with a growing demand for highly realistic, identity-preserving images even when only unpaired, unaligned datasets are available. We study unpaired face manipulation via adversarial learning, moving from autoencoder baselines to a robust, guided CycleGAN framework. While autoencoders capture coarse identity, they often miss fine details. Our approach integrates spectral normalization for stable training, identity- and perceptual-guided losses to preserve subject identity and high-level structure, and landmark-weighted cycle constraints to maintain facial geometry across pose and illumination changes. Experiments show that our adversarial trained CycleGAN improves realism (FID), perceptual quality (LPIPS), and identity preservation (ID-Sim) over autoencoders, with competitive cycle-reconstruction SSIM and practical inference times, which achieved high quality without paired datasets and approaching pix2pix on curated paired subsets. These results demonstrate that guided, spectrally normalized CycleGANs provide a practical path from autoencoders to robust unpaired face manipulation.

人脸生成无配对学习循环GAN对抗学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。