arXiv:2506.19288cs.CVcs.RO2025-06

首个面向水域的图像描述数据集,助力无人船理解复杂水道环境。

Da Yu: Towards USV-Based Image Captioning for Waterway Surveillance and Scene Understanding

  • 构建首个专用水域图像描述数据集WaterCaption,支持多区域长文本生成。
  • 提出轻量级模型Da Yu,在保持高效的同时提升长文本生成能力。
  • 适合水域监控、智能无人船场景理解的研究者与开发者使用。

自动化水道环境感知对无人水面艇(USV)理解周围环境并作出决策至关重要。现有感知模型多聚焦于实例级目标检测与分割,但受限于水道环境的复杂性,当前数据集和模型难以实现全局语义理解,制约了大规模监控与结构化日志生成。随着视觉-语言模型的发展,本文引入图像描述任务,构建首个专为水道环境设计的标注数据集WaterCaption,强调细粒度、多区域的长文本描述,为视觉地理理解与空间场景认知提供新方向。WaterCaption包含20.2万张图像-文本对,词汇量达180万。同时,提出可边缘部署的多模态大语言模型Da Yu,创新设计轻量级视觉-语言投影器Nano Transformer Adaptor(NTA),在计算效率与全局/局部特征建模能力间取得平衡,显著提升长文本生成性能。Da Yu在WaterCaption及其他多个描述基准上超越现有模型,实现性能与效率的最佳平衡。

原文摘要 · Abstract (English)

Automated waterway environment perception is crucial for enabling unmanned surface vessels (USVs) to understand their surroundings and make informed decisions. Most existing waterway perception models primarily focus on instance-level object perception paradigms (e.g., detection, segmentation). However, due to the complexity of waterway environments, current perception datasets and models fail to achieve global semantic understanding of waterways, limiting large-scale monitoring and structured log generation. With the advancement of vision-language models (VLMs), we leverage image captioning to introduce WaterCaption, the first captioning dataset specifically designed for waterway environments. WaterCaption focuses on fine-grained, multi-region long-text descriptions, providing a new research direction for visual geo-understanding and spatial scene cognition. Exactly, it includes 20.2k image-text pair data with 1.8 million vocabulary size. Additionally, we propose Da Yu, an edge-deployable multi-modal large language model for USVs, where we propose a novel vision-to-language projector called Nano Transformer Adaptor (NTA). NTA effectively balances computational efficiency with the capacity for both global and fine-grained local modeling of visual features, thereby significantly enhancing the model's ability to generate long-form textual outputs. Da Yu achieves an optimal balance between performance and efficiency, surpassing state-of-the-art models on WaterCaption and several other captioning benchmarks.

图像描述无人船多模态边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。