无需激光雷达,仅用摄像头实现多视角视频生成。
ReCamDriving: LiDAR-Free Camera-Controlled Video Synthesis for Novel Trajectories
- 用稠密3DGS渲染作为视觉几何引导,控制视频生成。
- 在11万对视频数据上训练,实现高结构一致性。
- 适合自动驾驶场景下多视角视频合成研究者使用。
多通道视频合成对自动驾驶至关重要。现有修复类方法常产生分布外伪影,相机控制类方法因稀疏激光雷达信号导致3D不一致。本文提出ReCamDriving,一种纯视觉框架,通过稠密且结构完整的3DGS渲染提供几何引导,实现相机可控生成。为防止模型过拟合于简单修复方案,采用两阶段渐进式训练:第一阶段用相机位姿进行粗略控制,第二阶段引入3DGS渲染实现精细视点与几何引导。此外,为对齐训练与推理的相机变换模式,提出基于3DGS的跨轨迹数据构建策略,使单次通行视频即可提供一致的横向轨迹监督。基于该策略,构建了包含约11万对平行轨迹视频的ParaDrive数据集。大量实验表明,ReCamDriving在相机可控性与结构一致性上达到当前最佳水平。
原文摘要 · Abstract (English)
Synthesizing multi-pass videos is important for autonomous driving. While current repair-based methods often struggle with out-of-distribution artifacts, camera-controlled methods often produce 3D-inconsistent results due to sparse LiDAR cues. We propose ReCamDriving, a purely vision-based framework that achieves camera-controlled generation by leveraging dense, structurally complete 3DGS renderings as geometric guidance. Specifically, to prevent the model from overfitting to a trivial repair solution when conditioning on 3DGS renderings, we adopt a two-stage progressive training paradigm: the first stage uses camera poses for coarse control, while the second stage incorporates 3DGS renderings for fine-grained viewpoint and geometric guidance. Furthermore, to align training and inference camera transformation patterns, we propose a 3DGS-based cross-trajectory data curation strategy, enabling consistent lateral-trajectory supervision from single-pass videos. Based on this strategy, we construct the ParaDrive dataset, containing approximately 110K parallel-trajectory video pairs. Extensive experiments demonstrate that ReCamDriving achieves state-of-the-art camera controllability and structural consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。