arXiv:2607.14710cs.CVcs.SY2026-07

用变分推理提升多摄像头自动驾驶鸟瞰图分割精度

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving

论文配图:Variational Inference for Bird's Eye View Segmentation in Autonomous Driving
图 1 · 摘自论文原文
  • 基于变分自编码器与归一化流生成多个鸟瞰图候选
  • 在nuScenes和OPV2V数据集上优于现有方法
  • 适合需要高精度环境感知的自动驾驶系统

鸟瞰图(BEV)已成为自动驾驶环境感知的关键技术,提供统一的空间表示。然而,如何有效融合多摄像头数据并在复杂驾驶环境中稳定运行仍是挑战。本文提出一种基于变换器的变分流变换网络TVB,将BEV分割问题置于变分推断框架下。通过后验BEV监督,网络隐式学习从多视角到统一标准BEV的映射。TVB以条件变分自编码器(CVAE)为骨干,生成多个BEV地图候选。通过引入归一化流增强生成地图的真实感,构建更复杂的概率分布。此外,设计了BEV注意力融合模块(BAF),利用注意力机制自适应整合多个候选地图。在nuScenes和OPV2V数据集上的实验表明,该方法在多视角BEV分割与车道环境感知任务中表现优异。

原文摘要 · Abstract (English)

The bird's eye view (BEV) has emerged as a pivotal approach for environmental perception in autonomous driving, providing a unified spatial representation for vehicles. Nevertheless, despite BEV's significance in addressing the challenges inherent to autonomous driving, effectively fusing data from multiple camera sensors and operating in complex external driving environments remains a considerable challenge. To mitigate this issue, we recast the BEV segmentation problem within a variational inference framework. In this paper, we propose a novel transformer-based variational flow transformation network for BEV segmentation, denoted as TVB. Our architecture implicitly learns the mapping from multiple camera views to a unified canonical BEV map during training by exploiting posterior BEV supervision. TVB employs a conditional variational auto encoder (CVAE) as its backbone and produces multiple BEV map candidates. To augment the realism of the generated BEV maps, we integrate normalizing flows into the map generation process, enabling the construction of more complex and expressive probability distributions. Furthermore, we design a BEV-attention fusion (BAF) module that harnesses attention mechanisms to adaptively integrate the multiple candidate BEV maps. Experimental results, evaluated on both the nuScenes and OPV2Vdatasets, demonstrate that our proposed method achieves superior performance in multi-camera view BEV segmentation and lane environment perception.

自动驾驶鸟瞰图变分推断多传感器融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。