用双分支注意力融合运动与语义信息,提升单目里程计鲁棒性。
MVOFormer: Flow-Semantic Transformer for Robust Monocular Visual Odometry

- 双分支编码器分别提取运动流与语义特征,区分静态与动态物体
- 迭代多模态解码器实现粗到精的位姿优化,动态抑制不可靠区域
- 无需微调即可跨数据集零样本泛化,适合复杂动态场景应用
单目视觉里程计(MVO)是自主导航与机器人定位的基础。现有学习型MVO方法常因缺乏可解释的互补特征或架构过于复杂,导致鲁棒性和跨域泛化能力受限。本文提出MVOFormer,一种新型变压器框架用于鲁棒单目视觉里程计。其采用流-语义双分支编码器,协同密集几何运动线索与以物体为中心的语义先验,明确区分静态结构与动态干扰物。这些表征由迭代多模态解码器融合,实现从粗到精的位姿细化,并动态抑制不可靠区域的关注度。大量实验表明,无需目标域微调,MVOFormer在TartanAir、KITTI、TUM-RGBD和ETH3D-SLAM等多个基准上均显著优于现有学习型帧间方法,展现出卓越的零样本泛化能力与鲁棒性。
原文摘要 · Abstract (English)
Monocular visual odometry (MVO) is foundational to autonomous navigation and robotic localization. However, existing learning-based MVO approaches often struggle with either a lack of interpretable, complementary features or overly complex multi-stage architectures. These limitations inherently restrict their robustness and cross-domain generalization. In this work, we propose MVOFormer, a novel transformer framework for robust monocular visual odometry. Our architecture features a Flow-Semantic Dual Branch Encoder that synergizes dense geometric motion cues with object-centric semantic priors, explicitly distinguishing static structures from dynamic distractors. These representations are then fused by an Iterative Multimodal Decoder, enabling coarse-to-fine pose refinement while dynamically suppressing attention on unreliable regions. Extensive evaluations demonstrate that, without any target-domain fine-tuning, MVOFormer achieves superior zero-shot generalization and robustness, significantly outperforming prior learning-based frame-to-frame methods across diverse benchmarks including TartanAir, KITTI, TUM-RGBD, and ETH3D-SLAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。