针对鱼眼相机的畸变问题,提出新框架实现更精准的鸟瞰图分割。
FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
- 设计抗畸变多尺度特征提取模块,保持尺度一致性。
- 引入不确定性感知注意力,提升多视角对齐可靠性。
- 动态平衡远近视野信息,确保时间连续性,适合自动驾驶场景。
作为自动驾驶的核心技术,鸟瞰图(BEV)分割在针孔相机上已取得显著进展。然而,将现有方法拓展至具有严重几何畸变、模糊多视角对应关系和不稳定时序动态的鱼眼相机仍具挑战,这些因素显著降低BEV性能。为此,我们提出FishBEV,一种专为鱼眼相机设计的新型BEV分割框架。该框架包含三项互补创新:(1)抗畸变多尺度特征提取(DRME)主干网络,在畸变环境下学习鲁棒特征并保持尺度一致性;(2)不确定性感知空间交叉注意力(U-SCA)机制,利用不确定性估计实现可靠的跨视角对齐;(3)距离感知时序自注意力(D-TSA)模块,自适应平衡近场细节与远场上下文,保障时序连贯性。在Synwoodscapes数据集上的大量实验表明,FishBEV在环绕式鱼眼相机的BEV分割任务中持续优于当前最优基线。
原文摘要 · Abstract (English)
As a cornerstone technique for autonomous driving, Bird's Eye View (BEV) segmentation has recently achieved remarkable progress with pinhole cameras. However, it is non-trivial to extend the existing methods to fisheye cameras with severe geometric distortion, ambiguous multi-view correspondences and unstable temporal dynamics, all of which significantly degrade BEV performance. To address these challenges, we propose FishBEV, a novel BEV segmentation framework specifically tailored for fisheye cameras. This framework introduces three complementary innovations, including a Distortion-Resilient Multi-scale Extraction (DRME) backbone that learns robust features under distortion while preserving scale consistency, an Uncertainty-aware Spatial Cross-Attention (U-SCA) mechanism that leverages uncertainty estimation for reliable cross-view alignment, a Distance-aware Temporal Self-Attention (D-TSA) module that adaptively balances near field details and far field context to ensure temporal coherence. Extensive experiments on the Synwoodscapes dataset demonstrate that FishBEV consistently outperforms SOTA baselines, regarding the performance evaluation of FishBEV on the surround-view fisheye BEV segmentation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。