用视觉大模型提升跨域场景重建精度,解决尺度不一致问题。
VFM-Recon: Unlocking Cross-Domain Scene-Level Neural Reconstruction with Scale-Aligned Foundation Priors
- 引入轻量级尺度对齐模块,恢复多视角尺度一致性。
- 在Tanks and Temples数据集上达到70.1的F1分数,远超对比方法。
- 适合需要跨域鲁棒性重建的三维视觉研究者。
单目视频驱动的场景级神经体素重建在严重域偏移下仍具挑战性。尽管视觉基础模型(VFMs)从大规模数据中学习到可迁移的通用先验,但其尺度模糊的预测与体素融合所需的尺度一致性不兼容。为此,我们提出VFM-Recon,首次将可迁移的VFM先验与场景级神经重建中的尺度一致性要求相结合。具体地,我们设计了一个轻量级尺度对齐阶段以恢复多视图尺度一致性,并通过轻量级任务专用适配器将预训练的VFM特征融入神经体素重建流程,在保持跨域鲁棒性的同时进行重建训练。模型在ScanNet训练集上训练,在同分布的ScanNet测试集以及跨分布的TUM RGB-D和Tanks and Temples数据集上评估。结果表明,该模型在所有数据集上均达到当前最优性能。尤其在具有挑战性的室外Tanks and Temples数据集上,重建网格的F1分数达70.1,显著优于最接近的对比方法VGGT(51.8)。
原文摘要 · Abstract (English)
Scene-level neural volumetric reconstruction from monocular videos remains challenging, especially under severe domain shifts. Although recent advances in vision foundation models (VFMs) provide transferable generalized priors learned from large-scale data, their scaleambiguous predictions are incompatible with the scale consistency required by volumetric fusion. To address this gap, we present VFMRecon, the first attempt to bridge transferable VFM priors with scaleconsistent requirements in scene-level neural reconstruction. Specifically, we first introduce a lightweight scale alignment stage that restores multiview scale coherence. We then integrate pretrained VFM features into the neural volumetric reconstruction pipeline via lightweight task-specific adapters, which are trained for reconstruction while preserving the crossdomain robustness of pretrained representations. We train our model on ScanNet train split and evaluate on both in-distribution ScanNet test split and out-of-distribution TUM RGB-D and Tanks and Temples datasets. The results demonstrate that our model achieves state-of-theart performance across all datasets domains. In particular, on the challenging outdoor Tanks and Temples dataset, our model achieves an F1 score of 70.1 in reconstructed mesh evaluation, substantially outperforming the closest competitor, VGGT, which only attains 51.8.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。