预训练视觉特征反而让少样本3D重建变差,挑战了主流认知。
Limitations of NERF with pre-trained Vision Features for Few-Shot 3D Reconstruction
- 用DINO特征增强NeRF,但效果不如基础NeRF
- 少样本下所有改进模型PSNR仅12.9–13.0,低于基线14.71
- 适合关注少样本3D重建可靠性与模型设计的研究者
神经辐射场(NeRF)已革新从稀疏图像集重建3D场景的方法。近期研究尝试引入预训练视觉特征(如DINO)以提升少样本重建性能,但在极端少样本场景下其有效性仍不明确。本文系统评估了基于DINO的NeRF模型,对比了基线NeRF、冻结的DINO特征、LoRA微调特征及多尺度特征融合方法。实验结果令人意外:所有DINO变体性能均劣于基线,PSNR值在12.9至13.0之间,而基线达14.71。这一反直觉现象表明,预训练视觉特征可能对少样本3D重建无益,甚至引入有害偏差。我们分析其原因包括特征-任务不匹配、有限数据下的过拟合以及融合难题。研究结果质疑领域常见假设,提示几何一致性优先的简单架构可能更适用于少样本场景。
原文摘要 · Abstract (English)
Neural Radiance Fields (NeRF) have revolutionized 3D scene reconstruction from sparse image collections. Recent work has explored integrating pre-trained vision features, particularly from DINO, to enhance few-shot reconstruction capabilities. However, the effectiveness of such approaches remains unclear, especially in extreme few-shot scenarios. In this paper, we present a systematic evaluation of DINO-enhanced NeRF models, comparing baseline NeRF, frozen DINO features, LoRA fine-tuned features, and multi-scale feature fusion. Surprisingly, our experiments reveal that all DINO variants perform worse than the baseline NeRF, achieving PSNR values around 12.9 to 13.0 compared to the baseline's 14.71. This counterintuitive result suggests that pre-trained vision features may not be beneficial for few-shot 3D reconstruction and may even introduce harmful biases. We analyze potential causes including feature-task mismatch, overfitting to limited data, and integration challenges. Our findings challenge common assumptions in the field and suggest that simpler architectures focusing on geometric consistency may be more effective for few-shot scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。