解决自动驾驶仿真中4D场景重建的物理一致性难题
Towards Physically Consistent 4D Scene Reconstruction for Closed-loop Autonomous Driving Simulation

- 提出分层训练方法OPG,分离时空参数以恢复空间可识别性
- 在多个数据集上实现稳定的新视角合成与更优的时序建模效果
- 适合需要高保真闭环仿真的自动驾驶研发团队使用
高保真街景重建对端到端自动驾驶仿真至关重要,其中新视角合成(NVS)与时变信息建模是闭环训练的基础能力。然而,现有3DGS及其4D扩展无法同时实现二者。我们建立信息几何诊断框架,揭示该问题源于时空参数间的信用分配困境:单源观测中视点与时间的确定性耦合导致低秩结构,引发静态视点依赖与动态时变成分之间的大量零空间模糊。时间信息压制空间线索,造成空间参数估计方差发散。为此,我们提出正交投影梯度(OPG)方法,先在初始阶段保障空间表示完整性,再将时间更新限制在空间零空间内,实现主动信用分配。同时引入基于物理先验的时序正则化策略,施加平滑约束以确保外观演化的一致性,从而保证闭环仿真中的物理一致性。大量实验表明,本方法不仅保持稳定的NVS能力,还在传统观测重现指标上表现更优,间接体现其对时序动态建模的提升。
原文摘要 · Abstract (English)
High-fidelity street scene reconstruction is pivotal for end-to-end autonomous driving simulation, where novel-view synthesis (NVS) and time-varying information modeling are two fundamental capabilities to facilitate closed-loop training. However, existing 3DGS methods and their 4D extensions fail to simultaneously achieve both. To bridge this gap, we establish an information-geometric diagnostic framework, revealing that this limitation stems from a credit assignment dilemma between spatial and temporal parameters. Specifically, the deterministic coupling between viewpoint and time in single-source observation creates a low-rank structure that induces massive null-space ambiguity between static view-dependent and dynamic time-varying components. Temporal information overshadows spatial cues, causing the estimation variance of spatial parameters to diverge. To address this issue, we propose Orthogonal Projected Gradient (OPG), a hierarchical training method designed to restore spatial identifiability. OPG prioritizes the integrity of spatial representations by securing them in an initial stage, then restricts temporal updates to the spatial null space, enabling proactive credit assignment. While OPG isolates temporal updates algebraically, Temporal Regularization Strategy is proposed to further refine the temporal solution space by imposing a smoothness constraint based on the physical prior of consistent appearance evolution, ensuring that the reconstructed scene remains physically consistent in closed-loop simulation. Extensive experiments demonstrate that our method not only maintains stable NVS capabilities but also demonstrates superior performance in traditional observation-reproducing metrics, which indirectly reflect the capability of modeling temporal dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。