arXiv:2512.05362cs.CVcs.LG2025-12

用深度学习快速判断视频是否适合做3D重建。

PoolNet: Deep Learning for 2D to 3D Video Process Validation

  • 设计端到端网络,从原始视频中判断能否进行3D重建
  • 在真实场景数据上准确率超90%,耗时仅为传统方法1/10
  • 适合需要批量筛选视频数据的3D重建研究者

从序列或非序列图像数据中提取结构光恢复(SfM)信息是一项耗时且计算成本高的任务。此外,大多数公开数据因相机位姿变化不足、遮挡或噪声问题而不适合处理。为此,我们提出PoolNet,一种用于帧级和场景级验证野外数据的通用深度学习框架。实验表明,该模型能有效区分适合与不适合进行SfM处理的场景,同时显著减少获取结构光恢复数据所需的时间,相比现有先进算法效率提升明显。

原文摘要 · Abstract (English)

Lifting Structure-from-Motion (SfM) information from sequential and non-sequential image data is a time-consuming and computationally expensive task. In addition to this, the majority of publicly available data is unfit for processing due to inadequate camera pose variation, obscuring scene elements, and noisy data. To solve this problem, we introduce PoolNet, a versatile deep learning framework for frame-level and scene-level validation of in-the-wild data. We demonstrate that our model successfully differentiates SfM ready scenes from those unfit for processing while significantly undercutting the amount of time state of the art algorithms take to obtain structure-from-motion data.

3D重建深度学习视频验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。