用跨任务注意力桥,让雷达相机联合做3D检测与分割
Radar-Camera BEV Multi-Task Learning with Cross-Task Attention Bridge for Joint 3D Detection and Segmentation

- 在鸟瞰图空间中,双向传递检测与分割特征
- 在nuScenes数据集上,7类分割性能提升,4类子集达51.0 mIoU-4
- 适合自动驾驶多任务感知系统研发者参考
鸟瞰图(BEV)表示已成为自动驾驶3D感知的主流范式,为检测与分割特征提供统一的空间坐标系。然而现有雷达-相机融合方法将这些任务孤立处理,错失了跨任务特征共享的机会:检测提供的物体级几何线索可增强分割精度,而分割提供的稠密道路布局上下文可约束检测定位。本文提出CTAB(Cross-Task Attention Bridge),一个在共享BEV空间中通过多尺度可变形注意力双向交换特征的模块。CTAB被集成进多任务框架,结合基于实例归一化的分割解码器和可学习的BEV上采样,以生成更精细的BEV表示。在nuScenes数据集上,CTAB在几乎不牺牲检测性能的前提下,使7类分割性能优于联合多任务基线;在4类子集(可行驶区域、人行横道、步行道、车辆)上,联合多任务模型达到51.0 mIoU-4的同时,仍保持竞争力的3D检测表现。
原文摘要 · Abstract (English)
Bird's-eye-view (BEV) representations are the dominant paradigm for 3D perception in autonomous driving, providing a unified spatial canvas where detection and segmentation features are geometrically registered to the same physical coordinate system. However, existing radar-camera fusion methods treat these tasks in isolation, missing the opportunity for cross-task feature sharing: object-level geometric cues from detection can sharpen segmentation, while dense road-layout context from segmentation can anchor detection. We propose \textbf{CTAB} (Cross-Task Attention Bridge), a bidirectional module that exchanges features between detection and segmentation branches via multi-scale deformable attention in shared BEV space. CTAB is integrated into a multi-task framework with an Instance Normalization-based segmentation decoder and learnable BEV upsampling to provide a more detailed BEV representation. On nuScenes, CTAB improves segmentation on 7 classes over the joint multi-task baseline at essentially neutral detection. On a 4-class subset (drivable area, pedestrian crossing, walkway, vehicle), our joint multi-task model achieves 51.0 mIoU-4 while simultaneously providing competitive 3D detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。