提出渐进式残差自回归框架,提升摄像头与雷达融合的鸟瞰图语义分割精度
RESAR-BEV: An Explainable Progressive Residual Autoregressive Approach for Camera-Radar Fusion in BEV Segmentation
- 分阶段残差自回归建模,通过双变压器结构实现从粗到细的可解释分割
- 在nuScenes上达到54.0% mIoU,推理速度达14.6 FPS,实现实时性与高精度兼顾
- 适用于自动驾驶中复杂场景下的多传感器融合,尤其适合长距与恶劣天气
鸟瞰图(BEV)语义分割为自动驾驶提供全面环境感知,但面临多模态错位和传感器噪声问题。本文提出RESAR-BEV,一种渐进式精炼框架,突破单步端到端方法:(1) 通过残差自回归学习实现渐进式精炼,利用驱动-变压器与修正-变压器级联架构,将BEV分割分解为可解释的粗到细阶段;(2) 采用贴近地面体素结合自适应高度偏移,并通过双路径体素特征编码(最大值+注意力池化)实现高效特征提取;(3) 采用离线真值分解与在线联合优化解耦监督,防止过拟合同时保证结构一致性。在nuScenes数据集上的实验表明,RESAR-BEV在7类关键驾驶场景中取得54.0% mIoU的最先进性能,且保持14.6 FPS的实时能力,对远距离感知和恶劣天气具有强鲁棒性。
原文摘要 · Abstract (English)
Bird's-Eye-View (BEV) semantic segmentation provides comprehensive environmental perception for autonomous driving but suffers multi-modal misalignment and sensor noise. We propose RESAR-BEV, a progressive refinement framework that advances beyond single-step end-to-end approaches: (1) progressive refinement through residual autoregressive learning that decomposes BEV segmentation into interpretable coarse-to-fine stages via our Drive-Transformer and Modifier-Transformer residual prediction cascaded architecture, (2) robust BEV representation combining ground-proximity voxels with adaptive height offsets and dual-path voxel feature encoding (max+attention pooling) for efficient feature extraction, and (3) decoupled supervision with offline Ground Truth decomposition and online joint optimization to prevent overfitting while ensuring structural coherence. Experiments on nuScenes demonstrate RESAR-BEV achieves state-of-the-art performance with 54.0% mIoU across 7 essential driving-scene categories while maintaining real-time capability at 14.6 FPS. The framework exhibits robustness in challenging scenarios of long-range perception and adverse weather conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。