用统一球面表示实现异构多相机系统快速鲁棒的视觉惯性里程计。
Sphere-VIO: Fast and Robust Visual-Inertial Odometry via Unified Spherical Representation for Heterogeneous Multi-Camera Systems

- 提出球面全景模型,统一映射多相机图像到共享球面空间,避免拼接延迟。
- 在公开数据集和自建数据集上实现高精度、稳定深度估计与实时性能。
- 适合资源受限设备部署,尤其适用于广角或全景多相机系统应用。
多相机视觉惯性里程计(VIO)通过扩展视场克服纯视觉系统的固有局限。然而,现有算法通常针对固定相机布局设计,缺乏对异构多相机系统的统一兼容性。由于缺乏统一的跨相机表示与匹配机制,当前方法难以在跨相机特征跟踪鲁棒性、深度估计稳定性与实时性能之间取得平衡。为此,本文提出 Sphere-VIO,一种基于轻量级滤波器框架、支持异构多相机系统的统一球面表示视觉惯性里程计。首先,提出统一球面全景模型(USPM),支持所有标准相机模型,无需顺序拼接即可实现多相机图像与共享球面空间间的双向快速映射,简化跨相机特征管理并提升三角化效率。其次,设计并行加速的深度引导半直接跟踪流水线——层级全向特征对齐(HOFA),引入全局球面约束以增强跨相机匹配鲁棒性,并将多相机深度观测融合至标准深度滤波器中实现稳定初始化。最后,开发适配多相机的扩展卡尔曼滤波后端(ESKF),采用球面方向残差与舒尔补边缘化策略,最小化计算开销,实现在资源受限设备上的高精度实时状态估计。大量实验在公开基准与自建全景数据集上表明,Sphere-VIO 在精度、鲁棒性、效率与跨相机通用性之间达到优异权衡。
原文摘要 · Abstract (English)
Multi-camera visual-inertial odometry (VIO) overcomes the inherent limitations of pure visual systems by expanding the field of view. However, existing algorithms are typically tailored for fixed camera setups and lack unified compatibility with heterogeneous multi-camera systems. Meanwhile, due to the absence of a unified cross-camera representation and association mechanism, current methods struggle to achieve a balance among robust cross-camera feature tracking, stable depth estimation, and reliable real-time performance. To address these issues, we present Sphere-VIO, a lightweight filter-based VIO framework with unified spherical representation for heterogeneous multi-camera systems. Specifically, we first propose a Unified Spherical Panorama Model (USPM) that supports all standard camera models and enables bidirectional fast mapping between multi-camera images and a shared spherical space without sequential stitching, simplifying cross-camera feature management and improving triangulation efficiency. Second, we design a parallel-accelerated depth-guided semi-direct tracking pipeline, namely Hierarchical Omnidirectional Feature Alignment (HOFA), with global spherical constraints for robust cross-camera matching, and fuse multi-camera depth observations into a standard depth filter for stable initialization. Finally, we develop a multi-camera-adapted ESKF backend that employs spherical bearing residuals and Schur complement marginalization to minimize computational overhead, enabling accurate real-time state estimation on resource-constrained devices. Extensive experiments on public benchmarks and a custom omnidirectional dataset show that Sphere-VIO achieves superior trade-offs between accuracy, robustness, efficiency, and cross-camera generality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。