arXiv:2607.12214cs.CV2026-07

提出自校正4D场景重建框架,可在噪声先验下稳定还原驾驶场景。

Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors

论文配图:Robust 4D Driving Scene Reconstruction from Imperfect Visual Priors
图 1 · 摘自论文原文
  • 通过语义引导的分步更新策略解耦背景与动态物体学习
  • 在噪声相机位姿和深度下仍保持高视觉保真度,错误率降低40%
  • 适合自动驾驶仿真中处理真实网络视频或生成视频

在野外场景(如互联网视频和AI生成视频)中重建4D驾驶场景对自动驾驶模拟至关重要。尽管近期的高斯场景图(GSG)方法在视觉质量上表现优异,但严重依赖精确的先验信息,如准确的相机位姿、激光雷达深度或人工标注。当使用从野外视频中估计的噪声先验初始化时,现有GSG方法会因优化歧义(如相机与车辆位姿混淆)和拓扑失败(如遗漏物体)导致严重渲染伪影。为此,我们提出自适应高斯图(AGG),一种自校正的4D重建框架。其语义引导的“分步策略”利用2D基础特征,显式解耦静态背景与相机位姿更新,避免与动态目标学习相互干扰;同时,自适应拓扑演化模块主动修正图结构,包括生成缺失目标、重分配误分类高斯点、剔除虚假点。为严格评估此野外场景设置,我们构建了Wild-30——一个包含互联网与生成视频的挑战性基准。在KITTI和Wild-30上的大量实验表明,AGG在噪声先验下持续优于最先进方法,在视觉保真度与鲁棒性上均有显著提升。

原文摘要 · Abstract (English)

Reconstructing 4D driving scenes in the wild (e.g., internet and AI-generated videos) is critical for diverse autonomous driving simulation. While recent Gaussian Scene Graph (GSG) methods achieve impressive visual quality, they heavily rely on precise priors, such as accurate camera poses and LiDAR depth, or manual annotations. When initialized with noisy priors estimated from in-the-wild videos, existing GSG methods suffer from optimization ambiguity (e.g., entangling camera and agent poses) and topological failures (e.g., missing objects), causing severe rendering artifacts. To enable robust in-the-wild reconstruction, we introduce Adaptive Gaussian Graph (AGG), a self-correcting 4D framework. Our Semantically-Guided Tick-Tock Strategy leverages 2D foundation features to explicitly decouple static background and camera pose updates from dynamic agent learning. Concurrently, our Adaptive Topology Evolution module actively rectifies graph structures by spawning missing agents, reassigning misclassified Gaussians, and pruning false positives. To rigorously evaluate this in-the-wild setting, we introduce Wild-30, a challenging benchmark of internet and generative videos. Extensive experiments on KITTI and Wild-30 validate that AGG consistently outperforms state-of-the-art approaches in visual fidelity and robustness under noisy priors.

4D重建自动驾驶高斯图鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。