用可学习的注意力机制选关键图像块,提升单目里程计在复杂环境下的精度和泛化能力。
DINO-VO: Learning Where to Focus for Enhanced State Estimation
- 引入可微分自适应图像块选择器,自动聚焦有效视觉区域。
- 在TartanAir等4个数据集上实现领先跟踪精度,尤其在合成与户外场景表现突出。
- 适合需要高鲁棒性单目定位的自动驾驶、机器人导航场景。
我们提出DINO Patch视觉里程计(DINO-VO),一种具备强场景泛化能力的端到端单目视觉里程计系统。现有视觉里程计常依赖启发式特征提取策略,在大规模室外环境中易导致精度和鲁棒性下降。DINO-VO通过在端到端流程中引入可微分的自适应图像块选择模块,提升所提取图像块的质量,并增强跨多样化数据集的泛化能力。此外,系统集成多任务特征提取模块与可微分束调整(BA)模块,利用逆深度先验,使模型能有效学习并利用外观与几何信息。该设计弥合了特征学习与状态估计之间的差距。在TartanAir、KITTI、Euroc和TUM数据集上的大量实验表明,DINO-VO在合成、室内和室外环境均展现出优异的泛化性能,达到当前最优的追踪精度。
原文摘要 · Abstract (English)
We present DINO Patch Visual Odometry (DINO-VO), an end-to-end monocular visual odometry system with strong scene generalization. Current Visual Odometry (VO) systems often rely on heuristic feature extraction strategies, which can degrade accuracy and robustness, particularly in large-scale outdoor environments. DINO-VO addresses these limitations by incorporating a differentiable adaptive patch selector into the end-to-end pipeline, improving the quality of extracted patches and enhancing generalization across diverse datasets. Additionally, our system integrates a multi-task feature extraction module with a differentiable bundle adjustment (BA) module that leverages inverse depth priors, enabling the system to learn and utilize appearance and geometric information effectively. This integration bridges the gap between feature learning and state estimation. Extensive experiments on the TartanAir, KITTI, Euroc, and TUM datasets demonstrate that DINO-VO exhibits strong generalization across synthetic, indoor, and outdoor environments, achieving state-of-the-art tracking accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。