arXiv:2411.19289cs.CV2024-11中稿 · IEEE/RSJ Internati…被引 1

解决动态视觉惯性里程计中提示不稳定导致的误差问题。

STAG-VIO: Stabilized Prompt-to-Geometry Interface for Robust Dynamic Visual--Inertial Odometry

  • 用自适应跟踪稳定提示,避免遮挡时检测抖动
  • 通过形态学优化建立保守安全边界,降低误判
  • 适合高动态场景下的机器人定位与导航

动态视觉惯性里程计(VIO)需可靠抑制运动干扰测量,但现有语义辅助方法依赖类别受限的分割器,在部分遮挡下性能下降。可提示的基础分割模型能实现无类别动态解析,但其在VIO中的效果高度依赖输入提示的时间稳定性——这一关键因素常被忽视。当原始检测生成的提示在遮挡下抖动或间断时,掩码会在帧间闪烁,导致几何估计失稳。我们提出STAG-VIO,将动态鲁棒性建模为感知到几何的接口稳定性问题。引入不确定性自适应多目标跟踪,将提示生成视为带有限噪声适应的状态估计,生成时间连贯的框提示。这些稳定提示驱动轻量级基础分割器,其掩码经几何导向的形态学精修以建立保守安全区域。一种约束预算感知的特征重分配策略,在动态区域占主导时仍能保留良好条件的静态测量。在VIODE和OpenLORIS-Scene数据集上的实验显示,持续优于现有最先进方法。消融实验表明,提示稳定是影响最大的单一组件,轨迹误差最多降低83%。

原文摘要 · Abstract (English)

Dynamic visual-inertial odometry (VIO) requires reliable suppression of motion-corrupted measurements, yet prior semantic-assisted approaches depend on category-limited segmenters and degrade under partial occlusion. Promptable foundation segmentation models offer category-agnostic dynamic parsing, but their effectiveness in VIO depends critically on the temporal stability of input prompts---a factor largely overlooked in existing pipelines. When prompts derived from raw detection are jittery or intermittent under occlusion, the resulting masks flicker across frames, destabilizing geometric estimation. We propose STAG-VIO, which formulates dynamic robustness as a perception-to-geometry interface stabilization problem. We introduce uncertainty-adaptive multi-object tracking that models prompt generation as state estimation with bounded noise adaptation, producing temporally coherent box prompts. These stabilized prompts drive a lightweight foundation segmenter whose masks undergo geometry-oriented morphological refinement to establish conservative safety margins. A constraint-budget-aware feature redistribution strategy preserves well-conditioned static measurements when dynamic regions dominate the view. Experiments on VIODE and OpenLORIS-Scene show consistent gains over state-of-the-art baselines. Ablation confirms that prompt stabilization is the single most impactful component, reducing trajectory error by up to 83%.

视觉惯性动态分割定位导航提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。