针对鱼眼相机与激光雷达融合的几何失真问题,提出新框架提升低重叠场景下的3D目标检测性能。
Geometry-Aware Fisheye-LiDAR Fusion for Robust 3D Object Detection in Low-Overlap Setups

- 在极坐标鸟瞰图中处理鱼眼图像特征,保留原始角度密度
- 通过注意力校正模块抑制边缘伪影,提升融合质量
- 首次实现鱼眼相机与激光雷达的有效融合,适合低成本自动驾驶系统
随着自动驾驶系统从高成本机器人出租车转向低成本物流场景,传感器配置越来越注重性价比。典型的稀疏视图方案采用双鱼眼相机搭配车顶激光雷达,带来严重几何挑战:极端径向畸变、重叠区域极少、球面投影与平面网格错位。现有鸟瞰图(BEV)融合方法通常早期将图像与点云统一到笛卡尔网格,导致广角鱼眼图像特征严重失真和信息丢失。为此,本文提出几何感知混合融合(GA-HF)框架,显式建模鱼眼几何与BEV特征畸变:通过畸变感知的升维-投影-射击(LSS)模块,将鱼眼特征映射至极坐标鸟瞰网格以保持原生角度密度;同时激光雷达特征在原始笛卡尔空间处理,确保边界框回归的度量精度。为桥接异构数据流,引入双注意力扭曲校正模块,在融合前对图像特征应用空间与通道注意力,明确抑制低质量边缘区域伪影,增强高质量语义线索。在KITTI-360、Dur360BEV和Fisheye3DOD三个基准上评估表明,该方法是首个探索激光雷达-鱼眼相机融合的方法。在KITTI-360上,相比笛卡尔基线模型,NDS提升4.2%;在Dur360BEV上,超越仅激光雷达与BEVFusion方法,且显著降低姿态误差;在Fisheye3DOD上,各项检测指标均领先于所有融合方法。
原文摘要 · Abstract (English)
As autonomous systems expand from capital-intensive robotaxis to cost-sensitive logistics, sensor configurations are increasingly optimized for coverage-per-cost. A prevalent sparse-view setup utilizes dual-fisheye cameras with a roof-mounted LiDAR, introducing severe geometric challenges: extreme radial distortion, minimal overlap, and misalignment between spherical projections and rectilinear grids. BEV fusion algorithms typically force image and point cloud modalities into unified Cartesian grids early in the pipeline, causing significant feature distortion and information loss for wide-view fisheye cameras. To address this, we propose a Geometry-Aware Hybrid Fusion (GA-HF) framework that explicitly accounts for fisheye geometry and BEV feature distortion, where fisheye features are lifted into a polar BEV grid via a Distortion-Aware Lift-Splat-Shoot (LSS) module to preserve native angular density, while LiDAR features are processed in native Cartesian space for metric fidelity of bounding box regression. To bridge these heterogeneous streams, we introduce a Dual-Attention Warping Correction module that applies spatial and channel attention to the warped camera features before fusion, explicitly suppressing artifacts in low-quality peripheral regions while enhancing high-quality semantic cues. GA-HF is evaluated on three benchmarks: KITTI-360, Dur360BEV, and Fisheye3DOD datasets. To the best of our knowledge, it is the first approach to explore LiDAR-fisheye camera fusion. On KITTI-360, GA-HF improves NDS by 4.2% over Cartesian baselines; on Dur360BEV, it surpasses both LiDAR-only and BEVFusion, while significantly reducing orientation error despite the geometric distortions; on Fisheye3DOD, it attains the highest detection score among all fusion methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。