arXiv:2412.19455cs.CV2024-12被引 1

用更少参数实现高质量动漫图像转换,且无需依赖低质配对数据。

NijiGAN: Transform What You See into Anime with Contrastive Semi-Supervised Learning and Neural Ordinary Differential Equations

  • 引入神经微分方程建模连续生成过程,提升转换精度。
  • 仅用一半参数量,FID达58.71,优于基线模型的60.32。
  • 通过伪配对数据训练,摆脱对低质量标注数据的依赖。

生成式AI已改变动画产业。现有图像到图像翻译模型多聚焦于将真实图像转为动漫风格,尤其关注无配对数据下的转换。Scenimefy利用对比学习,在半监督训练下缓解配对数据不足问题,但依赖经过微调的StyleGAN生成的动漫域配对数据,常导致数据质量低下;且其高参数量架构存在优化空间。本文提出NijiGAN,结合神经微分方程(NeuralODEs),在连续变换建模上具有优势。该模型仅使用Scenimefy生成的伪配对数据进行监督训练,避免了对低质量真实配对数据的依赖。实验表明,NijiGAN以半数参数量达到更高性能:在主观评分中,其均值意见得分(MOS)为2.192,高于AnimeGAN的2.160;在客观指标上,弗雷谢尔初始距离(FID)为58.71,优于Scenimefy的60.32。综合评估显示,该模型在多项指标上达到或超越现有先进水平。

原文摘要 · Abstract (English)

Generative AI has transformed the animation industry. Several models have been developed for image-to-image translation, particularly focusing on converting real-world images into anime through unpaired translation. Scenimefy, a notable approach utilizing contrastive learning, achieves high fidelity anime scene translation by addressing limited paired data through semi-supervised training. However, it faces limitations due to its reliance on paired data from a fine-tuned StyleGAN in the anime domain, often producing low-quality datasets. Additionally, Scenimefy's high parameter architecture presents opportunities for computational optimization. This research introduces NijiGAN, a novel model incorporating Neural Ordinary Differential Equations (NeuralODEs), which offer unique advantages in continuous transformation modeling compared to traditional residual networks. NijiGAN successfully transforms real-world scenes into high fidelity anime visuals using half of Scenimefy's parameters. It employs pseudo-paired data generated through Scenimefy for supervised training, eliminating dependence on low-quality paired data and improving the training process. Our comprehensive evaluation includes ablation studies, qualitative, and quantitative analysis comparing NijiGAN to similar models. The testing results demonstrate that NijiGAN produces higher-quality images compared to AnimeGAN, as evidenced by a Mean Opinion Score (MOS) of 2.192, it surpasses AnimeGAN's MOS of 2.160. Furthermore, our model achieved a Frechet Inception Distance (FID) score of 58.71, outperforming Scenimefy's FID score of 60.32. These results demonstrate that NijiGAN achieves competitive performance against existing state-of-the-arts, especially Scenimefy as the baseline model.

图像生成神经ODE动漫风格迁移半监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。