用虚拟位移自监督修复,让单目摄像头也能生成逼真新视角。
VisionNVS: Self-Supervised Inpainting for Novel View Synthesis under the Virtual-Shift Paradigm
- 将新视角合成转为自监督图像修复任务,利用原始图像做监督。
- 在无激光雷达条件下,几何精度和视觉质量优于基线方法。
- 适合大规模自动驾驶仿真,尤其擅长处理多相机系统偏差。
自动驾驶中的新视角合成(NVS)面临根本性挑战:模型在推理时需生成未见视角,但训练中缺乏这些视角的真实图像作为监督。本文提出VisionNVS,一种仅依赖摄像头的框架,将原本病态的外推问题重构为自监督图像修复任务。通过引入“虚拟位移”策略,利用单目深度代理模拟遮挡模式,并将其映射到原视图上,使原始记录图像可作为像素级精确的监督信号,有效消除先前方法中的领域差距。此外,我们提出伪3D接缝合成策略,利用相邻摄像头的视觉数据建模真实世界的光度差异与标定误差,提升空间一致性。实验表明,VisionNVS在无需激光雷达的情况下,显著优于依赖激光雷达的基线方法,在几何保真度与视觉质量方面表现更优,为可扩展的驾驶仿真提供了鲁棒解决方案。
原文摘要 · Abstract (English)
A fundamental bottleneck in Novel View Synthesis (NVS) for autonomous driving is the inherent supervision gap on novel trajectories: models are tasked with synthesizing unseen views during inference, yet lack ground truth images for these shifted poses during training. In this paper, we propose VisionNVS, a camera-only framework that fundamentally reformulates view synthesis from an ill-posed extrapolation problem into a self-supervised inpainting task. By introducing a ``Virtual-Shift'' strategy, we use monocular depth proxies to simulate occlusion patterns and map them onto the original view. This paradigm shift allows the use of raw, recorded images as pixel-perfect supervision, effectively eliminating the domain gap inherent in previous approaches. Furthermore, we address spatial consistency through a Pseudo-3D Seam Synthesis strategy, which integrates visual data from adjacent cameras during training to explicitly model real-world photometric discrepancies and calibration errors. Experiments demonstrate that VisionNVS achieves superior geometric fidelity and visual quality compared to LiDAR-dependent baselines, offering a robust solution for scalable driving simulation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。