arXiv:2512.17450cs.CVcs.LG2025-12被引 1

构建多模态海事数据集,提升夜间视觉识别鲁棒性

MULTIAQUA: A multimodal maritime dataset and robust training strategies for multimodal semantic segmentation

  • 融合可见光、热成像、激光雷达等多模态传感器数据
  • 仅用日间图像训练,模型在全黑环境下仍保持稳定性能
  • 适用于无人船在复杂光照条件下的场景理解任务

无人水面艇在运行中会遭遇多种复杂视觉环境,部分情况仅靠彩色相机难以解析。为拓展海事数据资源,本文提出新型多模态海事数据集MULTIAQUA(Multimodal Aquatic Dataset),包含同步、校准并标注的RGB、热成像、红外、激光雷达等多模态传感器数据。该数据集旨在支持监督学习方法,从多模态信息中提取有效特征,实现恶劣能见度条件下的高质量场景理解。为验证数据集价值,我们在具有挑战性的夜间测试集上评估了多种多模态方法,并提出更稳健的训练策略,使模型在近似全黑条件下仍能保持可靠表现。所提方法仅需日间图像即可训练出鲁棒深度神经网络,显著简化数据采集、标注与训练流程。

原文摘要 · Abstract (English)

Unmanned surface vehicles can encounter a number of varied visual circumstances during operation, some of which can be very difficult to interpret. While most cases can be solved only using color camera images, some weather and lighting conditions require additional information. To expand the available maritime data, we present a novel multimodal maritime dataset MULTIAQUA (Multimodal Aquatic Dataset). Our dataset contains synchronized, calibrated and annotated data captured by sensors of different modalities, such as RGB, thermal, IR, LIDAR, etc. The dataset is aimed at developing supervised methods that can extract useful information from these modalities in order to provide a high quality of scene interpretation regardless of potentially poor visibility conditions. To illustrate the benefits of the proposed dataset, we evaluate several multimodal methods on our difficult nighttime test set. We present training approaches that enable multimodal methods to be trained in a more robust way, thus enabling them to retain reliable performance even in near-complete darkness. Our approach allows for training a robust deep neural network only using daytime images, thus significantly simplifying data acquisition, annotation, and the training process.

多模态海事数据夜间识别深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。