arXiv:2607.23758cs.CV2026-07

无需逐场景优化,一键重建完整道路表面。

RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction

论文配图:RoadVGGT: Road-Structure-Aware Feed-Forward Road Surface Reconstruction
图 1 · 摘自论文原文
  • 基于几何基础模型,用多视角图像直接预测稠密高斯属性。
  • 通过置信度加权网格融合,实现大范围道路表面一致性重建。
  • 适合自动驾驶高精地图构建与仿真,支持多模态输出。

大规模道路表面重建支撑高精地图、自动驾驶感知、标注与仿真。现有专用优化方法虽能生成高质量道路表示,但通常需逐场景训练及依赖轨迹周边的场景特定覆盖设计,限制了新道路的可扩展重建。为此,我们提出 RoadVGGT,一种道路结构感知的前馈框架,无需测试时逐场景优化即可重建紧凑的高斯道路表面。RoadVGGT 利用几何基础模型结合多视角图像及给定的位姿与深度观测,通过学习的高斯头预测稠密像素对齐的高斯属性。为使这些稠密预测适用于大范围道路表面,我们将其对齐至一致的度量世界坐标系,并在道路对齐的 XY 平面上通过置信度加权网格融合冗余高斯点。类别感知分组与道路-人行道交界保护进一步在脆弱道路结构处控制融合。最终表示支持 RGB 和语义鸟瞰图、高程估计及新视角合成。RoadVGGT 消除了先前方法中对逐场景优化的需求,以紧凑高斯表示重建完整道路表面,并提升图像质量、语义映射与高程精度。大量实验验证了几何基础模型在可扩展前馈道路表面重建中的潜力。

原文摘要 · Abstract (English)

Large-scale road surface reconstruction supports high-definition mapping, autonomous-driving perception, annotation, and simulation. Existing road-specialized optimization methods can produce high-quality road representations, but they typically require per-scene training and scene-dependent coverage design around the driving trajectory, limiting scalable reconstruction over newly collected roads. To address these limitations, we introduce RoadVGGT, a road-structure-aware feed-forward framework that reconstructs compact Gaussian road surfaces without test-time per-scene optimization. RoadVGGT uses a geometric foundation model to exploit multi-view images together with provided pose and depth observations, and predicts dense pixel-aligned Gaussian attributes through a learned Gaussian head. To make these dense predictions usable for large road surfaces, we align them into a consistent metric world coordinate system and fuse redundant Gaussians on the road-aligned XY plane through confidence-weighted grid fusion. Category-aware grouping and road--sidewalk junction protection further control fusion around vulnerable road structures. The resulting representation supports RGB and semantic bird's-eye-view maps, elevation estimation, and novel view synthesis. RoadVGGT eliminates the need for per-scene optimization in prior methods, reconstructs complete road surfaces with a compact Gaussian representation, and improves image quality, semantic mapping, and elevation accuracy. Extensive experiments demonstrate the potential of geometric foundation models for scalable feed-forward road surface reconstruction.

道路重建高斯表示自动驾驶前馈框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。