让自动驾驶相机在没拍过的地方也能生成清晰视图。
Geo-EVS: Geometry-Conditioned Extrapolative View Synthesis for Autonomous Driving

- 用几何信息引导重投影,统一训练与推理路径。
- 在无密集真值时仍提升高角度和低覆盖区域的重建质量。
- 适合需要泛化到未采样场景的自动驾驶视觉系统。
外推式新视角合成可通过从异构传感器生成标准化虚拟视图,降低自动驾驶对相机阵列的依赖。现有方法在记录轨迹之外性能下降,因外推位姿提供的几何支持弱且缺乏密集目标视图监督。关键在于训练时显式暴露模型于轨迹外条件缺陷。我们提出Geo-EVS,一种在稀疏监督下的几何条件框架。该框架包含两个组件:几何感知重投影(GAR)利用微调的VGGT重建彩色点云,并将其重投影至观测与虚拟目标位姿,生成几何条件图;这一设计统一了训练与推理的重投影路径。人工伪影引导的潜在扩散(AGLD)在训练中注入重投影生成的伪影掩码,使模型学会在缺失支持下恢复结构。评估采用LiDAR-Projected Sparse-Reference(LPSR)协议,当密集外推视图真值不可用时亦可衡量性能。在Waymo数据集上,Geo-EVS显著提升稀疏视图合成质量与几何准确性,尤其在高俯仰角和低覆盖区域。同时改善下游3D检测表现。
原文摘要 · Abstract (English)
Extrapolative novel view synthesis can reduce camera-rig dependency in autonomous driving by generating standardized virtual views from heterogeneous sensors. Existing methods degrade outside recorded trajectories because extrapolated poses provide weak geometric support and no dense target-view supervision. The key is to explicitly expose the model to out-of-trajectory condition defects during training. We propose Geo-EVS, a geometry-conditioned framework under sparse supervision. Geo-EVS has two components. Geometry-Aware Reprojection (GAR) uses fine-tuned VGGT to reconstruct colored point clouds and reproject them to observed and virtual target poses, producing geometric condition maps. This design unifies the reprojection path between training and inference. Artifact-Guided Latent Diffusion (AGLD) injects reprojection-derived artifact masks during training so the model learns to recover structure under missing support. For evaluation, we use a LiDAR-Projected Sparse-Reference (LPSR) protocol when dense extrapolated-view ground truth is unavailable. On Waymo, Geo-EVS improves sparse-view synthesis quality and geometric accuracy, especially in high-angle and low-coverage settings. It also improves downstream 3D detection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。