arXiv:2503.00315cs.RO2025-03被引 2

用对抗生成框架迭代优化单目视觉惯性里程计,提升精度与可解释性。

XIRVIO: Critic-guided Iterative Refinement for Visual-Inertial Odometry with Explainable Adaptive Weighting

  • 基于Transformer的生成对抗网络,通过迭代修正姿态轨迹
  • 在KITTI数据集上达到与顶尖方法相当的定位误差
  • 自适应加权机制揭示传感器注意力分布,适合安全关键场景

我们提出XIRVIO,一种基于Transformer的生成对抗网络(GAN)框架,用于单目视觉惯性里程计(VIO)。以图像序列和6-DoF惯性测量为输入,XIRVIO的生成器通过迭代精修过程预测位姿轨迹,再由判别器评估并选择最优迭代结果。此外,自涌现的自适应传感器加权机制能揭示模型如何根据数据上下文关注不同传感器输入,增强了在安全关键应用中的可解释性。在KITTI数据集上的评估表明,XIRVIO在平移和旋转误差方面均达到现有学习型方法的先进水平。

原文摘要 · Abstract (English)

We introduce XIRVIO, a transformer-based Generative Adversarial Network (GAN) framework for monocular visual inertial odometry (VIO). By taking sequences of images and 6-DoF inertial measurements as inputs, XIRVIO's generator predicts pose trajectories through an iterative refinement process which are then evaluated by the critic to select the iteration with the optimised prediction. Additionally, the self-emergent adaptive sensor weighting reveals how XIRVIO attends to each sensory input based on contextual cues in the data, making it a promising approach for achieving explainability in safety-critical VIO applications. Evaluations on the KITTI dataset demonstrate that XIRVIO matches well-known state-of-the-art learning-based methods in terms of both translation and rotation errors.

视觉惯性生成对抗网络可解释性姿态估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。