arXiv:2410.15932cs.CV2024-10被引 3

提升单目俯视图分割精度,通过自校准注意力聚焦关键区域

Focus on BEV: Self-calibrated Cycle View Transformation for Monocular Birds-Eye-View Segmentation

  • 自校准跨视角变换模块,聚焦图像中与俯视图相关区域
  • 在nuScenes上达29.2% mIoU,Argoverse上达35.2% mIoU
  • 适合自动驾驶场景的单目俯视图语义理解任务

俯视图(BEV)分割旨在从视角图像建立空间映射并生成俯视语义图。近期研究因图像空间中缺乏俯视图感知特征而面临视角变换困难。为此,我们提出新的FocusBEV框架:(i) 自校准跨视角变换模块,抑制图像中无关俯视图的区域,聚焦于与俯视图相关的区域;(ii) 基于车辆运动的即插即用时空融合模块,利用记忆库挖掘俯视图空间中的时空结构一致性;(iii) 占位无关的IoU损失,缓解语义与位置不确定性。实验表明,该方法在两个主流基准上达到新最优性能:nuScenes上为29.2% mIoU,Argoverse上为35.2% mIoU。

原文摘要 · Abstract (English)

Birds-Eye-View (BEV) segmentation aims to establish a spatial mapping from the perspective view to the top view and estimate the semantic maps from monocular images. Recent studies have encountered difficulties in view transformation due to the disruption of BEV-agnostic features in image space. To tackle this issue, we propose a novel FocusBEV framework consisting of $(i)$ a self-calibrated cross view transformation module to suppress the BEV-agnostic image areas and focus on the BEV-relevant areas in the view transformation stage, $(ii)$ a plug-and-play ego-motion-based temporal fusion module to exploit the spatiotemporal structure consistency in BEV space with a memory bank, and $(iii)$ an occupancy-agnostic IoU loss to mitigate both semantic and positional uncertainties. Experimental evidence demonstrates that our approach achieves new state-of-the-art on two popular benchmarks,\ie, 29.2\% mIoU on nuScenes and 35.2\% mIoU on Argoverse.

俯视图分割单目感知自校准自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。