分离风格与内容,生成更真实的水下图像。
DISC-GAN: Disentangling Style and Content for Cluster-Specific Synthetic Underwater Image Generation
- 用聚类划分水体风格,分别训练以保留环境特征。
- 生成图像SSIM达0.9012,PSNR均值32.51 dB,FID仅13.37。
- 适合需要真实感水下合成数据的研究者。
本文提出一种新框架DISC-GAN,将风格-内容解耦与特定聚类训练结合,实现逼真的水下图像合成。水下成像受色度衰减和浑浊等光学现象影响,导致不同水域呈现显著风格差异(如色调变化、雾霾)。现有生成模型难以建模复杂多变的水下环境。为此,我们采用K-means聚类将数据集划分为风格特异的子域,使用独立编码器学习风格与内容的潜在表示,并通过自适应实例归一化(AdaIN)融合,解码生成最终图像。模型在各风格簇上独立训练,以保持领域特性。实验表明,该方法达到当前最优性能:结构相似性指数(SSIM)为0.9012,平均峰值信噪比(PSNR)为32.5118 dB,弗雷谢初始距离(FID)为13.3728。
原文摘要 · Abstract (English)
In this paper, we propose a novel framework, Disentangled Style-Content GAN (DISC-GAN), which integrates style-content disentanglement with a cluster-specific training strategy towards photorealistic underwater image synthesis. The quality of synthetic underwater images is challenged by optical due to phenomena such as color attenuation and turbidity. These phenomena are represented by distinct stylistic variations across different waterbodies, such as changes in tint and haze. While generative models are well-suited to capture complex patterns, they often lack the ability to model the non-uniform conditions of diverse underwater environments. To address these challenges, we employ K-means clustering to partition a dataset into style-specific domains. We use separate encoders to get latent spaces for style and content; we further integrate these latent representations via Adaptive Instance Normalization (AdaIN) and decode the result to produce the final synthetic image. The model is trained independently on each style cluster to preserve domain-specific characteristics. Our framework demonstrates state-of-the-art performance, obtaining a Structural Similarity Index (SSIM) of 0.9012, an average Peak Signal-to-Noise Ratio (PSNR) of 32.5118 dB, and a Frechet Inception Distance (FID) of 13.3728.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。