无需标定即可高效重建多车协同驾驶的动态场景
FRUC: Feedforward Dynamic Scene Reconstruction from Uncalibrated Collaborative Driving Views

- 用视觉地基几何变换器实现一次性无标定推理
- 在真实数据集上渲染质量与效率均达新最优
- 适合自动驾驶中跨车协同感知场景
我们提出FRUC,一种从前向3D高斯点阵框架,用于从非标定协同驾驶视角重建动态场景。现有方法常受限于严格的标定要求和缓慢的逐场景优化。本文将分布式多车网络视为时空无结构的本体多相机系统,核心挑战在于通过协作增强本体被遮挡区域的几何信息,同时不破坏本体已精确观测的可见几何,并保持重建效率。FRUC基于视觉地基几何变换器,实现灵活数量多车视角的一次性、无标定推理。为在非标定跨车错位下实现非破坏性几何补全,提出本体因果遮挡场,通过建模各车时空相关性显式推导遮挡演化作为潜在先验。基于此先验,将跨车融合形式化为零初始化注入的确定性残差去噪过程,将复杂的跨车融合转化为有界残差学习,实现鲁棒的协同盲区补全。在真实世界V2XReal和UrbanIng-V2X数据集上的大量评估表明,FRUC在动态协同驾驶环境的场景重建中达到新最先进水平,在渲染质量和效率上显著优于现有方法。
原文摘要 · Abstract (English)
We present FRUC, a feed-forward 3D Gaussian splatting framework for dynamic scene reconstruction from uncalibrated collaborative driving views. Existing multi-agent reconstruction frameworks are often hindered by rigid prerequisites, demanding precise spatial calibration and slow per-scene optimization. In this paper, we rethink this task by conceptualizing a distributed multi-vehicle network as a spatio-temporally unstructured ego-centric multi-camera system, where the core challenge lies in enhancing ego-centric occluded geometry through collaboration without degrading the ego's accurately observed visible geometry, while preserving reconstruction efficiency. For efficient reconstruction, FRUC is built upon a visual grounded geometric Transformer backbone to enable one-shot, calibration-free inference from a flexible number of multi-vehicle views. To achieve non-destructive geometric supplementation under uncalibrated cross-agent misalignment, FRUC first introduces an ego-centric causal occlusion field that explicitly derives occlusion evolution as latent priors by modeling agent-wise spatio-temporal correlations. Guided by these occlusion priors, it further formulates cross-agent integration as a deterministic residual denoising process via zero-initialized injection, turning challenging cross-agent fusion into bounded residual learning for robust collaborative blind-spot completion. Through extensive evaluations on the real-world V2XReal and UrbanIng-V2X datasets, FRUC is shown to be a new state-of-the-art for the scene reconstruction of dynamic collaborative driving environments, significantly outperforming existing methods in both rendering quality and efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。