构建系统化图像匹配基准,评估不同几何挑战下的方法表现。
RUBIK: A Structured Benchmark for Image Matching across Geometric Challenges
- 基于重叠度、缩放比和视角角设计33级难度标准。
- 最先进方法成功率达54.8%,仍难应对极端几何变化。
- 适合研究图像匹配鲁棒性与效率的开发者参考。
相机位姿估计在众多计算机视觉应用中至关重要,但现有基准难以揭示方法在不同几何挑战下的局限性。我们提出RUBIK,一个新型基准,系统评估图像匹配方法在明确几何难度等级下的表现。基于重叠度、缩放比和视角角三个互补指标,我们将nuScenes数据集中的16.5K张图像对划分为33个难度级别。对14种方法的全面评估显示,尽管无检测器方法表现最优(成功率>47%),但其计算开销显著高于基于检测器的方法(150-600ms vs. 40-70ms)。即使表现最佳的方法也仅在54.8%的图像对上成功,表明在低重叠、大尺度差异与极端视角变化共现的挑战场景中仍有巨大改进空间。该基准将公开发布。
原文摘要 · Abstract (English)
Camera pose estimation is crucial for many computer vision applications, yet existing benchmarks offer limited insight into method limitations across different geometric challenges. We introduce RUBIK, a novel benchmark that systematically evaluates image matching methods across well-defined geometric difficulty levels. Using three complementary criteria - overlap, scale ratio, and viewpoint angle - we organize 16.5K image pairs from nuScenes into 33 difficulty levels. Our comprehensive evaluation of 14 methods reveals that while recent detector-free approaches achieve the best performance (>47% success rate), they come with significant computational overhead compared to detector-based methods (150-600ms vs. 40-70ms). Even the best performing method succeeds on only 54.8% of the pairs, highlighting substantial room for improvement, particularly in challenging scenarios combining low overlap, large scale differences, and extreme viewpoint changes. Benchmark will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。