arXiv:2410.08612cs.CVcs.AI2024-10被引 4

用双扩散模型和GPT提示生成逼真多样的声呐图像。

Synth-SONAR: Sonar Image Synthesis with Enhanced Diversity and Realism via Dual Diffusion Models and GPT Prompting

  • 双级文本条件扩散模型分步生成粗略到精细的声呐图像。
  • 生成数据集规模大,真实感与多样性显著提升。
  • 首次应用GPT提示生成声呐图像,适合海洋探测研究者。

声呐图像合成对水下勘探、海洋生物学和国防应用至关重要。传统方法依赖昂贵且耗时的声呐传感器数据采集,影响数据质量和多样性。为此,本文提出新型声呐图像合成框架Synth-SONAR,融合扩散模型与GPT提示技术。核心创新包括:其一,结合生成式AI风格注入与公开的真实/模拟数据,构建目前最大的声呐数据集;其二,采用双级文本条件声呐扩散模型,分步生成粗粒度与细粒度图像,提升质量与多样性;其三,利用视觉语言模型(VLMs)和GPT提示中的高级语义信息,实现从文本提示生成高保真声呐图像。该方法首次在声呐成像中应用GPT提示,显著提升合成数据的真实感与多样性,达到当前最佳性能。

原文摘要 · Abstract (English)

Sonar image synthesis is crucial for advancing applications in underwater exploration, marine biology, and defence. Traditional methods often rely on extensive and costly data collection using sonar sensors, jeopardizing data quality and diversity. To overcome these limitations, this study proposes a new sonar image synthesis framework, Synth-SONAR leveraging diffusion models and GPT prompting. The key novelties of Synth-SONAR are threefold: First, by integrating Generative AI-based style injection techniques along with publicly available real/simulated data, thereby producing one of the largest sonar data corpus for sonar research. Second, a dual text-conditioning sonar diffusion model hierarchy synthesizes coarse and fine-grained sonar images with enhanced quality and diversity. Third, high-level (coarse) and low-level (detailed) text-based sonar generation methods leverage advanced semantic information available in visual language models (VLMs) and GPT-prompting. During inference, the method generates diverse and realistic sonar images from textual prompts, bridging the gap between textual descriptions and sonar image generation. This marks the application of GPT-prompting in sonar imagery for the first time, to the best of our knowledge. Synth-SONAR achieves state-of-the-art results in producing high-quality synthetic sonar datasets, significantly enhancing their diversity and realism.

声呐生成扩散模型GPT提示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。