融合雷达与视觉信息,提升水面目标检测精度
physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection
- 引入物理先验的雷达编码器,增强点云可靠性与散射特征
- 通过雷达引导的交互式融合,实现跨模态语义对齐,提升长距检测能力
- 适合水下/水面无人艇感知、复杂海况下的多传感器融合研究
针对无人水面艇(USV)在波浪杂波、镜面反射和远距离弱外观线索下的水面目标检测难题,本文提出PhysFusion——一种基于Transformer的双流雷达-视觉融合框架。该框架包含:(1) 物理信息雷达编码器(PIR Encoder),结合雷达散射截面映射与质量门控机制,将点级雷达属性转化为紧凑散射先验并预测可靠性;(2) 雷达引导的交互式融合模块(RIFM),在查询级别实现语义丰富的雷达特征与多尺度视觉特征融合,雷达分支采用基于点的局部流与基于Transformer的全局流,使用散射感知自注意力(SASA)建模;(3) 时序查询聚合模块(TQA),在短时窗口内聚合帧间融合查询以获得时序一致性表征。在WaterScenes和FLOW数据集上验证,当雷达历史长度T=5时,PhysFusion在WaterScenes上达到59.7% mAP50:95和90.3% mAP50,仅需5.6M参数和12.5G FLOPs;在FLOW上雷达+相机设置下达94.8% mAP50和46.2% mAP50:95。消融实验量化了PIR Encoder、SASA全局推理与RIFM的贡献。
原文摘要 · Abstract (English)
Detecting water-surface targets for Unmanned Surface Vehicles (USVs) is challenging due to wave clutter, specular reflections, and weak appearance cues in long-range observations. Although 4D millimeter-wave radar complements cameras under degraded illumination, maritime radar point clouds are sparse and intermittent, with reflectivity attributes exhibiting heavy-tailed variations under scattering and multipath, making conventional fusion designs struggle to exploit radar cues effectively. We propose PhysFusion, a physics-informed radar-image detection framework for water-surface perception. The framework integrates: (1) a Physics-Informed Radar Encoder (PIR Encoder) with an RCS Mapper and Quality Gate, transforming per-point radar attributes into compact scattering priors and predicting point-wise reliability for robust feature learning under clutter; (2) a Radar-guided Interactive Fusion Module (RIFM) performing query-level radar-image fusion between semantically enriched radar features and multi-scale visual features, with the radar branch modeled by a dual-stream backbone including a point-based local stream and a transformer-based global stream using Scattering-Aware Self-Attention (SASA); and (3) a Temporal Query Aggregation module (TQA) aggregating frame-wise fused queries over a short temporal window for temporally consistent representations. Experiments on WaterScenes and FLOW demonstrate that PhysFusion achieves 59.7% mAP50:95 and 90.3% mAP50 on WaterScenes (T=5 radar history) using 5.6M parameters and 12.5G FLOPs, and reaches 94.8% mAP50 and 46.2% mAP50:95 on FLOW under radar+camera setting. Ablation studies quantify the contributions of PIR Encoder, SASA-based global reasoning, and RIFM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。