用视觉语言一致性提升3D模型训练效果,无需额外标注。
VLRC: Vision-Language Reprojection Consistency as a scalable signal for better feed-forward 3D pretraining

- 利用冻结的视觉语言模型作为多视角语义监督信号。
- 在室内和室外数据集上提升3D重建精度与零样本语义分割性能。
- 适合做无标注场景下3D预训练的模型改进。
前馈式3D模型通常依赖昂贵的几何监督或自监督的光度目标,二者提供的学习信号均不完整。本文提出视觉-语言重投影一致性(VLRC),一种可扩展的辅助目标,利用冻结的视觉-语言表征作为语义多视角监督。给定一个预测的3D重建结果,VLRC将密集的视觉-语言特征跨视角重投影,并强制对应图像位置间的特征一致性,无需额外3D标注。该目标能无缝融入自监督单目重建及已预训练的监督模型在无标签适应阶段。通过将几何结构与语言引导的特征对齐,VLRC不仅改善深度与相机估计,还实现更连贯的多视角语义融合,支持开放词汇的3D场景理解。在室内与室外基准测试中,持续提升3D重建准确率与零样本开放词汇3D语义分割性能。
原文摘要 · Abstract (English)
Feed-forward 3D models are commonly trained using either expensive geometric supervision or self-supervised photometric objectives, both of which provide incomplete learning signals. We introduce Vision-Language Reprojection Consistency (VLRC), a scalable auxiliary objective that exploits frozen vision-language representations as semantic multi-view supervision. Given a predicted 3D reconstruction, VLRC reprojects dense vision-language features across views and enforces feature consistency between corresponding image locations, requiring no additional 3D annotations. The objective integrates seamlessly with both self-supervised monocular reconstruction and supervised-pretrained feed-forward 3D models during unlabeled adaptation. By aligning geometry with language-grounded features, VLRC not only improves depth and camera estimation but also enables more coherent multi-view semantic fusion for open-vocabulary 3D scene understanding. Experiments on indoor and outdoor benchmarks demonstrate consistent gains in 3D reconstruction accuracy and zero-shot open-vocabulary 3D semantic segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。