用跨域学习解决鱼眼图像生成数据少难题,让自动驾驶感知更全面。
MIVIFI: Bridging Perspective and Fisheye Domains for Training Multi-View Fisheye Image Generation Models

- 利用全景投影桥接鱼眼与标准视角数据,实现跨域训练
- 可在有限鱼眼数据上生成高保真多视角鱼眼图像,支持气象光照变化
- 适合自动驾驶场景生成、仿真数据增强等研究者使用
实现360°覆盖对自动驾驶视觉感知至关重要。鱼眼相机仅需两个传感器即可提供全向视野,但现有多视角鱼眼数据集稀缺,罕见场景通常需昂贵的3D仿真生成,制约模型训练。尽管生成模型在标准透视图像中取得显著进展,其在广角畸变图像中的应用仍属空白。本文首次提出基于体素语义表示的多视角鱼眼图像生成任务,并提出两种方法:首先将SyntheOcc架构适配至鱼眼数据(SyntheOcc-FE),但受限于鱼眼数据稀缺,泛化能力不足;随后提出MIVIFI方法,通过等距柱状投影实现跨域学习,结合KITTI-360鱼眼图像与nuScenes多视角标准图像,实现高质量场景内容操控。该框架可对语义占据输入进行结构修改,实现特定目标物的增删,并生成原数据集中缺失的多种气象与光照条件。定量与定性实验表明,所提方法能生成鲁棒且逼真的多视角鱼眼图像,凸显跨域策略在缓解数据稀缺问题上的优势。
原文摘要 · Abstract (English)
Achieving 360° coverage is critical for the visual perception systems of autonomous vehicles. Fisheye cameras offer a cost-effective solution by enabling full surround coverage with as few as two sensors. However, existing multi-view fisheye datasets are limited, and synthesizing rare corner cases typically requires computationally expensive 3D simulations, hindering the training. While generative models have achieved significant success in standard perspective imagery, their application to wide-angle distortion remains unexplored. In this work, we formally introduce the novel problem of multi-view fisheye image generation conditioned on volumetric semantic representations and present two distinct methods. We first propose SyntheOcc-FE, which adapts the SyntheOcc architecture to fisheye data. While effective, this method is constrained by the scarcity of fisheye datasets, which limits its generalization. To overcome these limitations, we propose our second method, MIVIFI (multi-view fisheye), which leverages cross-domain learning with Equirectangular Projections. By bridging the gap between dataset domains using KITTI-360 fisheye images alongside nuScenes multi-view standard images, our approach enables high-fidelity manipulation of scene content. This framework enables the structural modification of semantic occupancy inputs to introduce or eliminate specific actors and facilitates the rendering of diverse meteorological conditions and illumination scenarios absent in the limited fisheye datasets. Quantitative and qualitative experiments demonstrate that our methods achieve robust photorealistic multi-view fisheye image generation and highlight the specific advantages of our cross-domain strategy for handling data scarcity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。