arXiv:2412.13188cs.CV2024-12CVPR被引 53

用激光点云控制生成逼真街景视频,视角变化更自由。

StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models

论文配图:StreetCrafter: Street View Synthesis with Controllable Video Diffusion Models
图 1 · 摘自论文原文
  • 以激光点云为条件,控制视频扩散模型生成街景。
  • 在Waymo和PandaSet上实现更大范围的高质量视图合成。
  • 支持像素级编辑,适合自动驾驶场景生成与渲染。

本文针对车辆传感器数据下的逼真视图合成问题提出StreetCrafter,一种可控视频扩散模型。该模型利用激光雷达点云渲染作为像素级条件,充分挖掘生成先验以实现新视角合成,同时保持精确相机控制。像素级点云条件使目标场景可进行精准像素级修改。此外,模型生成先验可有效融入动态场景表示,实现实时渲染。在Waymo Open Dataset与PandaSet上的实验表明,该方法能灵活控制视角变化,显著扩展视图合成区域,性能优于现有方法。

原文摘要 · Abstract (English)

This paper aims to tackle the problem of photorealistic view synthesis from vehicle sensor data. Recent advancements in neural scene representation have achieved notable success in rendering high-quality autonomous driving scenes, but the performance significantly degrades as the viewpoint deviates from the training trajectory. To mitigate this problem, we introduce StreetCrafter, a novel controllable video diffusion model that utilizes LiDAR point cloud renderings as pixel-level conditions, which fully exploits the generative prior for novel view synthesis, while preserving precise camera control. Moreover, the utilization of pixel-level LiDAR conditions allows us to make accurate pixel-level edits to target scenes. In addition, the generative prior of StreetCrafter can be effectively incorporated into dynamic scene representations to achieve real-time rendering. Experiments on Waymo Open Dataset and PandaSet demonstrate that our model enables flexible control over viewpoint changes, enlarging the view synthesis regions for satisfying rendering, which outperforms existing methods.

视频生成扩散模型街景合成激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。