融合多传感器信息,让车载设备生成更真实的街景新视角视频。
Effective Multi-sensor Conditioning for Street-view Novel-view Synthesis

- 用多传感器信号联合建模,提升新视角生成精度。
- 在稀疏激光雷达下表现超越现有方法,等效于10-100倍密集点云。
- 适合自动驾驶、虚拟导航等需要高保真场景重演的场景。
现代车载平台配备丰富的传感器套件,包括激光雷达、校准的多相机阵列和精确的自车位姿,理论上可为从新视角重渲染驾驶场景提供强信号。近年来的工作利用视频扩散模型,基于生成先验从稀疏车辆观测中合成合理的新型视图。然而,现有方法仅利用了部分信号,当目标轨迹偏离原始路径时,生成质量显著下降。我们认为这本质上是一个多传感器融合问题:稀疏激光雷达投影提供准确但不完整的度量几何,环视影像提供密集外观但无度量深度,相机位姿则跨视角连接两者。我们提出StreetNVS,一种视频扩散框架,通过基于相对射线级位置编码的参考增强相机注意力模块,联合条件于三种信号。我们设计了两阶段课程训练策略,逐步暴露模型于越来越稀疏的激光雷达数据。在Waymo Open Dataset上,StreetNVS在稀疏激光雷达条件下显著优于当前最优基线,性能相当于依赖10-100倍密集点云的方法。我们还展示了沿极端非轨迹路径(如抬升、变道、后退、旋转)生成连贯视频的能力。
原文摘要 · Abstract (English)
Modern vehicle platforms are equipped with a rich sensor suite, including LiDAR, calibrated multi-camera rigs, and accurate ego-motion, that in principle offers strong signal for re-rendering a driving scene from novel viewpoints. A growing line of recent work leverages video diffusion models for this task, using their generative priors to synthesize plausible novel views from sparse vehicle observations. In practice, however, existing methods exploit only a fragment of this signal, and their quality tends to degrade as the target trajectory departs from the recorded driving path. We argue that this is fundamentally a multi-sensor fusion problem: sparse LiDAR reprojections supply accurate but incomplete metric geometry, surround-view reference imagery supplies dense appearance but no metric depth, and camera poses tie the two together across views. We introduce StreetNVS, a video diffusion framework that jointly conditions on all three signals through a Reference-Enhanced Camera Attention module based on a relative ray-level positional encoding. We develop a two-stage curriculum training strategy that gradually exposes the model to increasingly sparse LiDAR. On the Waymo Open Dataset, StreetNVS substantially outperforms state-of-the-art baselines under sparse LiDAR conditioning, matches methods that rely on 10-100 times denser point clouds. We further show capabilities of synthesizing coherent videos along extreme out-of-trajectory paths such as elevation, lane-shift, pullback, and rotation. Our website: https://streetnvs.github.io
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。