用可变形注意力提升关键点检测鲁棒性,适配大视角变化场景。
RDD: Robust Feature Detector and Descriptor using Deformable Transformer
- 基于可变形Transformer捕捉全局上下文与几何不变性
- 在稀疏匹配任务中超越所有现有方法,支持半稠密匹配
- 构建空地跨域新基准,验证复杂视角下的性能优势
作为结构光恢复与SLAM的核心步骤,面对显著视角变化等挑战场景时,关键点的鲁棒检测与描述仍未解决。尽管近期工作强调局部特征对几何变换建模的重要性,但这些方法未能学习长程关系中的视觉线索。本文提出新型鲁棒关键点检测器/描述子RDD,利用可变形Transformer通过可变形自注意力机制捕捉全局上下文与几何不变性。具体而言,可变形注意力聚焦于关键位置,有效降低搜索空间复杂度并建模几何不变性。此外,我们新增了一个空地数据集用于训练,同时使用标准MegaDepth数据集。所提方法在稀疏匹配任务中优于所有前沿方法,并具备半稠密匹配能力。为确保全面评估,我们引入两个挑战性基准:一个强调大幅视角与尺度变化,另一个为空地跨域基准——这一设置近年来在不同高度的3D重建中日益流行。
原文摘要 · Abstract (English)
As a core step in structure-from-motion and SLAM, robust feature detection and description under challenging scenarios such as significant viewpoint changes remain unresolved despite their ubiquity. While recent works have identified the importance of local features in modeling geometric transformations, these methods fail to learn the visual cues present in long-range relationships. We present Robust Deformable Detector (RDD), a novel and robust keypoint detector/descriptor leveraging the deformable transformer, which captures global context and geometric invariance through deformable self-attention mechanisms. Specifically, we observed that deformable attention focuses on key locations, effectively reducing the search space complexity and modeling the geometric invariance. Furthermore, we collected an Air-to-Ground dataset for training in addition to the standard MegaDepth dataset. Our proposed method outperforms all state-of-the-art keypoint detection/description methods in sparse matching tasks and is also capable of semi-dense matching. To ensure comprehensive evaluation, we introduce two challenging benchmarks: one emphasizing large viewpoint and scale variations, and the other being an Air-to-Ground benchmark -- an evaluation setting that has recently gaining popularity for 3D reconstruction across different altitudes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。