融合视觉与雷达数据,提升水面目标检测鲁棒性
WS-DETR: Robust Water Surface Object Detection through Vision-Radar Fusion with Detection Transformer
- 引入多尺度边缘增强与分层特征聚合模块
- 在复杂光照下仍保持最佳检测性能
- 适合水上无人艇等强干扰场景应用
针对复杂水环境下的水面无人艇(USVs)目标检测难题,本文提出一种鲁棒的视觉-雷达融合模型WS-DETR。针对水面物体边缘模糊、尺度多样等问题,设计了多尺度边缘信息融合(MSEII)模块增强边缘感知,并引入分层特征聚合(HiFA)提升多尺度检测能力。采用自移动点表示实现连续卷积与残差连接,高效提取不规则点云特征。为缓解跨模态特征冲突,提出自适应特征交互融合(AFIF)模块,通过几何对齐与语义融合整合视觉与雷达特征。在WaterScenes数据集上的大量实验表明,WS-DETR达到当前最优性能,且在恶劣天气和光照条件下依然保持优势。
原文摘要 · Abstract (English)
Robust object detection for Unmanned Surface Vehicles (USVs) in complex water environments is essential for reliable navigation and operation. Specifically, water surface object detection faces challenges from blurred edges and diverse object scales. Although vision-radar fusion offers a feasible solution, existing approaches suffer from cross-modal feature conflicts, which negatively affect model robustness. To address this problem, we propose a robust vision-radar fusion model WS-DETR. In particular, we first introduce a Multi-Scale Edge Information Integration (MSEII) module to enhance edge perception and a Hierarchical Feature Aggregator (HiFA) to boost multi-scale object detection in the encoder. Then, we adopt self-moving point representations for continuous convolution and residual connection to efficiently extract irregular features under the scenarios of irregular point cloud data. To further mitigate cross-modal conflicts, an Adaptive Feature Interactive Fusion (AFIF) module is introduced to integrate visual and radar features through geometric alignment and semantic fusion. Extensive experiments on the WaterScenes dataset demonstrate that WS-DETR achieves state-of-the-art (SOTA) performance, maintaining its superiority even under adverse weather and lighting conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。