arXiv:2505.14296cs.CVeess.IV2025-05

用对比学习提升水下图像生成真实感,深度信息让画面更逼真。

Towards Generating Realistic Underwater Images

  • 用对比学习替代循环一致性,增强空间结构保留
  • 加入深度信息后生成图像FID最低,真实感最强
  • 适合需要高真实感水下图像的科研与影视制作

本文研究了对比学习与生成对抗网络在从均匀光照合成图像生成逼真水下图像中的应用。基于VAROS数据集,评估了图像翻译模型的表现。使用弗雷谢特初始距离(FID)和结构相似性指数(SSIM)衡量感知质量与结构保真度之间的权衡。成对方法中,pix2pix因成对监督和PatchGAN判别器取得最佳FID;自编码器模型则获得最高SSIM,表明结构保真度更高但图像更模糊。无配对方法中,CycleGAN凭借循环一致性损失实现有竞争力的FID;CUT通过对比学习替代循环一致性,获得更高SSIM,说明空间相似性更好。值得注意的是,将深度信息融入CUT后,整体FID降至最低,表明深度线索显著提升真实感;但SSIM略有下降,提示深度感知学习可能引入结构变化。

原文摘要 · Abstract (English)

This paper explores the use of contrastive learning and generative adversarial networks for generating realistic underwater images from synthetic images with uniform lighting. We investigate the performance of image translation models for generating realistic underwater images using the VAROS dataset. Two key evaluation metrics, Fréchet Inception Distance (FID) and Structural Similarity Index Measure (SSIM), provide insights into the trade-offs between perceptual quality and structural preservation. For paired image translation, pix2pix achieves the best FID scores due to its paired supervision and PatchGAN discriminator, while the autoencoder model attains the highest SSIM, suggesting better structural fidelity despite producing blurrier outputs. Among unpaired methods, CycleGAN achieves a competitive FID score by leveraging cycle-consistency loss, whereas CUT, which replaces cycle-consistency with contrastive learning, attains higher SSIM, indicating improved spatial similarity retention. Notably, incorporating depth information into CUT results in the lowest overall FID score, demonstrating that depth cues enhance realism. However, the slight decrease in SSIM suggests that depth-aware learning may introduce structural variations.

图像生成对比学习水下图像生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。