arXiv:2501.02913cs.CV2025-01被引 4

用稀疏激光点云引导扩散模型,生成真实驾驶场景的连贯新视角图像。

Pointmap-Conditioned Diffusion for Consistent Novel View Synthesis

  • 用点图(点云栅格化)作为条件信号,引导2D扩散模型生成图像。
  • 在真实驾驶数据上实现高质量、视角一致的新视图合成,支持稀疏点云输入。
  • 可适配不同密度点图,且能用于构建3D高斯泼溅等后续表示。

生成外推视角仍是难题,尤其在城市驾驶场景中,可用数据仅限于有限的RGB图像和稀疏的LiDAR点。为此,我们提出PointmapDiff框架,利用预训练的2D扩散模型进行新视图合成。该方法以点图(即3D场景坐标的栅格化表示)作为条件信号,从参考图像中提取几何与光照先验,指导图像生成过程。通过引入参考注意力层和点图特征的ControlNet,PointmapDiff能够在不同视角间生成准确且一致的结果,并保持几何真实性。在真实驾驶数据上的实验表明,该方法对点图条件信号具有高度灵活性(如密集深度图或稀疏LiDAR点),并可用于蒸馏生成3D高斯泼溅等3D表示,进一步提升视图外推能力。

原文摘要 · Abstract (English)

Synthesizing extrapolated views remains a difficult task, especially in urban driving scenes, where the only reliable sources of data are limited RGB captures and sparse LiDAR points. To address this problem, we present PointmapDiff, a framework for novel view synthesis that utilizes pre-trained 2D diffusion models. Our method leverages point maps (i.e., rasterized 3D scene coordinates) as a conditioning signal, capturing geometric and photometric priors from the reference images to guide the image generation process. With the proposed reference attention layers and ControlNet for point map features, PointmapDiff can generate accurate and consistent results across varying viewpoints while respecting geometric fidelity. Experiments on real-life driving data demonstrate that our method achieves high-quality generation with flexibility over point map conditioning signals (e.g., dense depth map or even sparse LiDAR points) and can be used to distill to 3D representations such as 3D Gaussian Splatting for improving view extrapolation.

新视角合成扩散模型点云引导自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。