arXiv:2604.22714cs.CV2026-04被引 3

解决互联网照片长尾分布下的稀疏重建难题

Long-tail Internet photo reconstruction

论文配图:Long-tail Internet photo reconstruction
图 1 · 摘自论文原文
  • 用真实地标采样构建稀疏图像集模拟长尾场景
  • 在极端稀疏下仍可获得鲁棒3D重建,精度提升显著
  • 适合处理低密度、对称重复场景,兼容主流数据集

互联网照片集合呈现极长尾分布:少数著名地标被密集拍摄并易于3D重建,而大多数真实场景仅拥有稀疏、噪声大且不均匀的图像,超出传统与学习型3D方法的能力。我们认为攻克这一长尾情形是3D基础模型的下一前沿。尽管从稀疏场景获取可靠真值3D监督困难,我们发现可通过从已良好重建的互联网地标中采样稀疏子集来有效模拟。为此,我们提出MegaDepth-X,一个包含干净、稠密深度的大规模3D重建数据集,并设计了一种策略,用于采样模仿长尾场景相机分布的训练图像集。使用这些组件微调3D基础模型,可在极端稀疏条件下实现稳健重建,同时提升对称与重复场景的可靠性,并保持在标准稠密3D基准数据集上的泛化能力。

原文摘要 · Abstract (English)

Internet photo collections exhibit an extremely long-tailed distribution: a few famous landmarks are densely photographed and easily reconstructed in 3D, while most real-world sites are represented with sparse, noisy, uneven imagery beyond the capabilities of both classical and learned 3D methods. We believe that tackling this long-tail regime represents one of the next frontiers for 3D foundation models. Although reliable ground-truth 3D supervision from sparse scenes is challenging to acquire, we observe that it can be effectively simulated by sampling sparse subsets from well-reconstructed Internet landmarks. To this end, we introduce MegaDepth-X, a large dataset of 3D reconstructions with clean, dense depth, together with a strategy for sampling sets of training images that mimic camera distributions in long-tail scenes. Finetuning 3D foundation models with these components yields robust reconstructions under extreme sparsity, and also enables more reliable reconstruction in symmetric and repetitive scenes, while preserving generalization to standard, dense 3D benchmark datasets.

3D重建长尾分布稀疏图像基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。