arXiv:2503.21771cs.CV2025-03CVPR被引 7

用文本生成水下图像与密集标注,解决数据稀缺问题。

A Unified Image-Dense Annotation Generation Model for Underwater Scenes

论文配图:A Unified Image-Dense Annotation Generation Model for Underwater Scenes
图 1 · 摘自论文原文
  • 单模型同时生成水下图像和多种密集标注。
  • 引入布局共享与时序自适应归一化,保证图像与标注一致。
  • 可生成大规模数据集,提升水下感知模型性能。

水下密集预测(如深度估计、语义分割)对全面理解水下场景至关重要,但高质量、大规模带密集标注的水下数据集因环境复杂和采集成本高而极为稀缺。本文提出一种统一的文本到图像与密集标注生成方法(TIDE),仅需文本输入即可同步生成逼真的水下图像和多种高度一致的密集标注。通过在单一模型中统一图像与标注生成任务,引入隐式布局共享机制(ILS)和跨模态交互方法时间自适应归一化(TAN),联合优化图像与密集标注的一致性。利用TIDE构建大规模水下数据集,验证了该方法在水下密集预测任务中的有效性。实验表明,该方法显著提升了现有水下密集预测模型的性能,并缓解了密集标注数据稀缺的问题。我们希望此方法能为其他领域数据稀缺问题提供新思路。代码已开源。

原文摘要 · Abstract (English)

Underwater dense prediction, especially depth estimation and semantic segmentation, is crucial for gaining a comprehensive understanding of underwater scenes. Nevertheless, high-quality and large-scale underwater datasets with dense annotations remain scarce because of the complex environment and the exorbitant data collection costs. This paper proposes a unified Text-to-Image and DEnse annotation generation method (TIDE) for underwater scenes. It relies solely on text as input to simultaneously generate realistic underwater images and multiple highly consistent dense annotations. Specifically, we unify the generation of text-to-image and text-to-dense annotations within a single model. The Implicit Layout Sharing mechanism (ILS) and cross-modal interaction method called Time Adaptive Normalization (TAN) are introduced to jointly optimize the consistency between image and dense annotations. We synthesize a large-scale underwater dataset using TIDE to validate the effectiveness of our method in underwater dense prediction tasks. The results demonstrate that our method effectively improves the performance of existing underwater dense prediction models and mitigates the scarcity of underwater data with dense annotations. We hope our method can offer new perspectives on alleviating data scarcity issues in other fields. The code is available at https://github.com/HongkLin/TIDE

水下视觉图像生成密集标注数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。