arXiv:2508.15930cs.CV2025-08被引 3

用视觉语言模型提升遥感图像船舶检测的语义感知能力

Semantic-Aware Ship Detection with Vision-Language Integration

  • 融合视觉语言模型与多尺度滑动窗口策略
  • 构建专用数据集ShipSem-VL捕捉船舶细粒度属性
  • 在复杂场景下显著提升检测精度,适合遥感应用研究者

遥感图像中的船舶检测在海洋活动监控、航运物流和环境研究中具有重要意义。现有方法常难以捕捉细粒度语义信息,限制了其在复杂场景下的效果。为此,我们提出一种结合视觉语言模型(VLMs)与多尺度自适应滑动窗口策略的新检测框架。为实现语义感知船舶检测(SASD),我们构建了专用的ShipSem-VL视觉语言数据集,以捕捉船舶的细粒度属性。通过三个明确的任务评估,全面分析了该框架的性能,从多个角度验证了其在推进SASD方面的有效性。

原文摘要 · Abstract (English)

Ship detection in remote sensing imagery is a critical task with wide-ranging applications, such as maritime activity monitoring, shipping logistics, and environmental studies. However, existing methods often struggle to capture fine-grained semantic information, limiting their effectiveness in complex scenarios. To address these challenges, we propose a novel detection framework that combines Vision-Language Models (VLMs) with a multi-scale adaptive sliding window strategy. To facilitate Semantic-Aware Ship Detection (SASD), we introduce ShipSem-VL, a specialized Vision-Language dataset designed to capture fine-grained ship attributes. We evaluate our framework through three well-defined tasks, providing a comprehensive analysis of its performance and demonstrating its effectiveness in advancing SASD from multiple perspectives.

船舶检测视觉语言模型遥感图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。