用视觉与对称先验,从单侧图像重建带真实尺度的完整车辆3D模型。
Leveraging Visual and Geometric Priors for Metric-scale and Complete Vehicle Gaussian Reconstruction from Limited Views

- 结合视觉基础模型与对称性先验,初始化并优化3D高斯表示。
- 在公开数据集上,完整度和几何精度均显著优于现有方法。
- 适合需要高保真车辆资产的自动驾驶仿真与长尾场景生成。
高保真车辆资产对可控交通场景生成至关重要,尤其适用于合成稀有且安全关键的长尾场景。然而,从真实环境车载图像中重建可复用的车辆表示仍面临两大挑战:其一,图像到3D生成方法通常无法保证可靠的真实尺度;其二,车载摄像头通常仅观测车辆一侧,导致传统多视图重建在未观测区域不完整。为此,我们提出一种前馈式车辆资产重建方法,利用两种互补先验,基于稀疏单侧观测重建车辆的3D高斯表示。为实现真实尺度重建,首先使用视觉基础模型作为视觉先验初始化高斯,再通过可学习编码器-解码器模块估计高斯属性。提出一种感知对称性的克隆策略,在高斯空间直接补全未观测侧,利用车辆的双边结构作为几何先验。在公开数据集上的实验表明,所提方法在车辆资产完整性和几何准确性方面均显著优于现有方法。
原文摘要 · Abstract (English)
High-fidelity vehicle assets are essential for controllable traffic scene generation, particularly for synthesizing rare and safety-critical long-tail scenarios. However, reconstructing a reusable vehicle representation from in-the-wild onboard images remains challenging for two reasons. First, image-to-3D generation methods generally produce models without reliable metric scale. Second, onboard cameras usually observe only one side of a target vehicle, making conventional multi-view reconstruction incomplete on unobserved regions. To solve these problems, we propose a feed-forward vehicle asset reconstruction method, which leverages two complementary priors to reconstruct 3D Gaussian representations for vehicles using sparse one-sided observations. To achieve metric-scale reconstruction, a visual foundation model is first utilized to serve as a visual prior for Gaussian initialization. The Gaussian attributes are then estimated by a learnable encoder-decoder module. A symmetry-aware cloning strategy is presented to complete the unobserved side directly in Gaussian space, which exploits the bilateral structure of vehicles as a geometric prior. Experiments on the public dataset demonstrate that the proposed method significantly outperforms existing approaches in both vehicle asset completeness and geometric accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。