arXiv:2507.17089cs.CVcs.RO2025-07被引 1

提出IONext模型,用CNN提升惯性里程计的精度与泛化能力。

IONext: Unlocking the Next Era of Inertial Odometry

  • 设计双翼自适应动态混合模块,捕捉全局与局部运动特征。
  • 在6个公开数据集上,平均ATE降低10%,RTE降低12%。
  • 适合需要高精度定位的自动驾驶与机器人场景。

研究人员越来越多地采用基于Transformer的模型进行惯性里程计建模。尽管Transformer擅长捕捉长程依赖关系,但其对局部细微运动变化敏感度不足且缺乏固有归纳偏置,常导致定位精度和泛化能力受限。近期研究表明,将大卷积核与Transformer-inspired结构引入CNN可有效扩展感受野,提升全局运动感知能力。受此启发,我们提出一种新型基于CNN的模块——双翼自适应动态混合器(DADM),可自适应地从动态输入中捕获全局运动模式与局部细微运动特征,并根据输入动态生成选择性权重,实现高效多尺度特征融合。为进一步改善时序建模,我们引入时空门控单元(STGU),在时域中选择性提取代表性且任务相关的运动特征,解决现有CNN方法在时序建模上的局限性。基于DADM与STGU,我们构建了新的基于CNN的惯性里程计骨干网络IONext。在六个公开数据集上的大量实验表明,IONext始终优于当前最先进(SOTA)的Transformer与CNN方法。例如,在RNIN数据集上,相比代表性模型iMOT,IONext将平均绝对轨迹误差(ATE)降低10%,平均相对轨迹误差(RTE)降低12%。

原文摘要 · Abstract (English)

Researchers have increasingly adopted Transformer-based models for inertial odometry. While Transformers excel at modeling long-range dependencies, their limited sensitivity to local, fine-grained motion variations and lack of inherent inductive biases often hinder localization accuracy and generalization. Recent studies have shown that incorporating large-kernel convolutions and Transformer-inspired architectural designs into CNN can effectively expand the receptive field, thereby improving global motion perception. Motivated by these insights, we propose a novel CNN-based module called the Dual-wing Adaptive Dynamic Mixer (DADM), which adaptively captures both global motion patterns and local, fine-grained motion features from dynamic inputs. This module dynamically generates selective weights based on the input, enabling efficient multi-scale feature aggregation. To further improve temporal modeling, we introduce the Spatio-Temporal Gating Unit (STGU), which selectively extracts representative and task-relevant motion features in the temporal domain. This unit addresses the limitations of temporal modeling observed in existing CNN approaches. Built upon DADM and STGU, we present a new CNN-based inertial odometry backbone, named Next Era of Inertial Odometry (IONext). Extensive experiments on six public datasets demonstrate that IONext consistently outperforms state-of-the-art (SOTA) Transformer- and CNN-based methods. For instance, on the RNIN dataset, IONext reduces the average ATE by 10% and the average RTE by 12% compared to the representative model iMOT.

惯性导航深度学习姿态估计CNN

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。