arXiv:2605.17327cs.ROcs.AI2026-05

用3D生成模型实现无需特征点的单目惯性系统快速初始化

Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using a Feed-Forward 3D Model

论文配图:Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using a Feed-Forward 3D Model
图 1 · 摘自论文原文
  • 直接用前馈3D模型生成点云,跳过传统特征追踪
  • 成功初始化率超90%,所需数据时长低于1.2秒
  • 在弱纹理环境仍稳定,适合复杂场景应用

单目视觉惯性导航系统(VINS)的快速可靠初始化至关重要,它决定了后续状态估计的起始条件。尽管进展不断,现有方法大多依赖视觉特征对应,需3-4秒传感数据才能成功初始化,限制了其适用性和效率。随着可直接从图像预测点云的前馈3D模型出现,我们重新审视了视觉惯性初始化问题。本文提出一种无特征初始化框架,利用前馈3D模型生成的未缩放点云,从而无需视觉特征跟踪与估计。该设计显著降低系统复杂度,提升初始化可靠性。在公开数据集上的实验表明,所提方法成功率达90%以上,且所需数据时长通常低于1.2秒。进一步在自采涵盖室内外多种场景的数据集上验证,证明其在视觉退化环境中表现稳健,优于现有方法。代码与数据集见:https://github.com/Yuantai-Z/FF-VIO-Init。

原文摘要 · Abstract (English)

Fast and reliable initialization is critical for monocular visual-inertial navigation systems (VINS), as it establishes the starting conditions for subsequent state estimation. Despite steady progress, most existing methods heavily rely on visual feature correspondences and require 3-4 seconds of sensory data for successful initialization, which limits their applicability and efficiency. With the advent of feed-forward 3D models that can directly predict point clouds from images, we revisit the visual-inertial initialization problem from a concise perspective. In this work, we propose a feature-free initialization framework that leverages up-to-scale point clouds predicted by a feed-forward 3D model, thereby obviating the need for visual feature tracking and estimation. This design substantially reduces system complexity and improves the reliability of initialization. Experiments on public datasets demonstrate that the proposed feature-free initialization method achieves the highest success rate, exceeding 90%, and significantly reduces the data duration required for successful initialization, typically to under 1.2 s. We further validate our method on a self-collected dataset covering various indoor and outdoor scenarios, demonstrating robust performance, particularly in visually degraded environments where existing methods often fail. The code and dataset are available at https://github.com/Yuantai-Z/FF-VIO-Init.

视觉惯性3D生成初始化无特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。