arXiv:2512.23983cs.CV2025-12

仅用图像和相机位姿实现驾驶场景的4D重建与视图外推。

DriveExplorer: Images-Only Decoupled 4D Reconstruction with Progressive Restoration for Driving View Extrapolation

  • 通过融合静态与动态点云,构建统一4D表示。
  • 利用扩散模型逐步修复并优化视角外推结果。
  • 无需激光雷达或标注,适合真实自动驾驶部署。

本文提出一种高效的自动驾驶场景视图外推方法。现有方法依赖激光雷达点云、3D边界框或车道标注等先验信息,需昂贵传感器或人工标注,限制实际应用。本文仅使用图像和可选相机位姿,首先估计全局静态点云与每帧动态点云,并融合为统一表示;随后采用可变形4D高斯框架进行场景重建。初始训练的4D高斯模型生成低质量伪图像,用于训练视频扩散模型;之后,通过扩散模型迭代优化逐步偏移的高斯渲染结果,将增强结果回传作为4DGS的训练数据,持续更新直至达到目标视角。相比基线方法,本方案在新视角上生成更高质量图像。

原文摘要 · Abstract (English)

This paper presents an effective solution for view extrapolation in autonomous driving scenarios. Recent approaches focus on generating shifted novel view images from given viewpoints using diffusion models. However, these methods heavily rely on priors such as LiDAR point clouds, 3D bounding boxes, and lane annotations, which demand expensive sensors or labor-intensive labeling, limiting applicability in real-world deployment. In this work, with only images and optional camera poses, we first estimate a global static point cloud and per-frame dynamic point clouds, fusing them into a unified representation. We then employ a deformable 4D Gaussian framework to reconstruct the scene. The initially trained 4D Gaussian model renders degraded and pseudo-images to train a video diffusion model. Subsequently, progressively shifted Gaussian renderings are iteratively refined by the diffusion model,and the enhanced results are incorporated back as training data for 4DGS. This process continues until extrapolation reaches the target viewpoints. Compared with baselines, our method produces higher-quality images at novel extrapolated viewpoints.

4D重建视图外推扩散模型自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。