无需真实几何数据即可评估和调优三维重建系统。
Look Ma, No Ground Truth! Ground-Truth-Free Tuning of Structure from Motion and Visual SLAM
- 通过对比原始与噪声图像的输出差异,实现无真值评估。
- 评估结果与传统方法高度相关,支持超参数自动调优。
- 适合大规模数据、自监督学习及实时场景下的系统优化。
结构从运动(SfM)和视觉即时定位与地图构建(VSLAM)系统的开发与调优依赖高质量的几何真值,但这类数据获取成本高、耗时长,且在多数实际场景中不可得。为此,本文提出一种全新的无真值(GTF)评估方法,通过分析输入图像在原始与添加噪声版本下的输出敏感性,实现无需几何真值的评估。该方法与传统基于真值的基准表现高度一致,并支持无真值条件下的超参数调优。摆脱对真值的依赖,使系统能利用更广泛的多源数据集,推动自监督与在线调优的发展,有望在数据驱动层面带来类似生成式AI的突破。
原文摘要 · Abstract (English)
Evaluation is critical to both developing and tuning Structure from Motion (SfM) and Visual SLAM (VSLAM) systems, but is universally reliant on high-quality geometric ground truth -- a resource that is not only costly and time-intensive but, in many cases, entirely unobtainable. This dependency on ground truth restricts SfM and SLAM applications across diverse environments and limits scalability to real-world scenarios. In this work, we propose a novel ground-truth-free (GTF) evaluation methodology that eliminates the need for geometric ground truth, instead using sensitivity estimation via sampling from both original and noisy versions of input images. Our approach shows strong correlation with traditional ground-truth-based benchmarks and supports GTF hyperparameter tuning. Removing the need for ground truth opens up new opportunities to leverage a much larger number of dataset sources, and for self-supervised and online tuning, with the potential for a data-driven breakthrough analogous to what has occurred in generative AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。