arXiv:2606.31147cs.CV2026-06

WaterGen可独立控制水下场景与介质效果,生成逼真多样数据集。

WaterGen: Decoupling Scene and Medium in Underwater Image Generation

论文配图:WaterGen: Decoupling Scene and Medium in Underwater Image Generation
图 1 · 摘自论文原文
  • 分两阶段生成:先建模场景内容,再添加水体退化效果
  • 合成数据提升水下修复与分割任务性能,效果优于基线
  • 适合需要可控水下图像数据的研究者,如视觉重建与算法验证

水下计算机视觉任务受限于大规模、多样化训练数据的匮乏。本文提出WaterGen,一种可独立控制场景与水体介质条件的大规模真实水下图像生成方法。该方法将水下图像生成解耦为两个因素:真实多样的场景内容(图像中有什么)和精确可控的水体介质效应(水如何影响图像)。现有方法要么控制性差,要么真实性不足,或无法独立建模介质。核心洞见是:在潜在扩散框架中可分离场景生成与介质建模,实现场景多样性与水体效果精准控制的兼顾。具体分为两步:首先,使用无退化的水下图像微调潜在扩散U-Net,学习生成不含介质退化的场景潜在嵌入;其次,将物理准确的介质退化合成建模为对这些潜在嵌入的条件解码过程。该解耦设计使模型能生成多样场景并完全控制水下外观。我们利用WaterGen构建了大规模合成水下数据集,其场景结构多样且水体效应准确,附带伪标签。实验表明,合成数据在水下修复和语义分割任务上持续提升下游性能。

原文摘要 · Abstract (English)

Underwater computer vision tasks, such as detection, restoration, and segmentation, are limited by the scarcity of large-scale and diverse training data. We introduce WaterGen, a method for generating large-scale, realistic, and diverse underwater images that provides independent control of the scene and water medium conditions. Our approach treats underwater image generation as the decoupled control of two factors: realistic and diverse scene content (what is in the image), and accurate and controllable water medium effects (what the water does to the image). Existing methods generally achieve only part of this objective: they either provide controllability with limited realism or diversity, or generate realistic scenes without accurately and independently modeling water-medium effects. Our key insight, that allows us to avoid this compromise, is that scene generation and medium modeling can be decoupled within a latent diffusion framework, enabling diverse scene generation together with accurate and controllable underwater appearance. To do this, we decompose underwater image synthesis into two stages. First, we fine-tune the latent diffusion U-Net using degradation-free underwater images so that it learns to generate diverse and realistic latent embeddings of underwater scene content without medium-induced degradation. Second, we formulate the physically accurate medium degradation synthesis as a conditional decoding process applied to these latent embeddings. This decoupled design allows our model to generate diverse scenes with full control of underwater appearance. We leverage WaterGen to build large-scale synthetic underwater datasets that are diverse in scene structures and accurate in water effects and pseudo-labels. We demonstrate that our synthetic data consistently improve downstream performance in underwater restoration and semantic segmentation.

图像生成水下视觉扩散模型数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。