arXiv:2506.07013cs.CV2025-06被引 7

UNO统一自监督单目里程计,跨平台无须调参即可精准定位。

UNO: Unified Self-Supervised Monocular Odometry for Platform-Agnostic Deployment

  • 用专家混合策略区分不同运动模式,自动选择最优解算器。
  • 在KITTI、EuRoC-MAV、TUM-RGBD三数据集上均达顶尖性能。
  • 适合自动驾驶、无人机、手持设备等多场景部署,无需重新训练。

本文提出UNO,一种统一的单目视觉里程计框架,可在多种环境、平台和运动模式下实现鲁棒且自适应的位姿估计。不同于依赖特定部署调优或预设运动先验的传统方法,本方法能有效泛化至真实世界多种场景,包括自动驾驶车辆、空中无人机、移动机器人及手持设备。为此,我们引入专家混合策略进行局部状态估计,采用多个专用解码器分别处理不同类别的自身运动模式。同时,设计全可微分的Gumbel-Softmax模块,构建帧间关联图,选择最优专家解码器并剔除错误估计。这些信息随后输入统一后端,结合预训练的尺度无关深度先验与轻量级捆绑调整,以保证几何一致性。我们在三个主流基准数据集上进行充分评估:KITTI(室外/自动驾驶)、EuRoC-MAV(室内/无人机)、TUM-RGBD(室内/手持),结果表明性能达到当前最优水平。

原文摘要 · Abstract (English)

This work presents UNO, a unified monocular visual odometry framework that enables robust and adaptable pose estimation across diverse environments, platforms, and motion patterns. Unlike traditional methods that rely on deployment-specific tuning or predefined motion priors, our approach generalizes effectively across a wide range of real-world scenarios, including autonomous vehicles, aerial drones, mobile robots, and handheld devices. To this end, we introduce a Mixture-of-Experts strategy for local state estimation, with several specialized decoders that each handle a distinct class of ego-motion patterns. Moreover, we introduce a fully differentiable Gumbel-Softmax module that constructs a robust inter-frame correlation graph, selects the optimal expert decoder, and prunes erroneous estimates. These cues are then fed into a unified back-end that combines pre-trained, scale-independent depth priors with a lightweight bundling adjustment to enforce geometric consistency. We extensively evaluate our method on three major benchmark datasets: KITTI (outdoor/autonomous driving), EuRoC-MAV (indoor/aerial drones), and TUM-RGBD (indoor/handheld), demonstrating state-of-the-art performance.

单目里程计自监督多平台部署专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。