arXiv:2606.25699cs.RO2026-06

解决多传感器定位中因环境退化导致的精度下降问题。

SA-LIVO: Efficient LiDAR-Inertial-Visual Odometry with Subspace-Aware Degeneracy Handling

论文配图:SA-LIVO: Efficient LiDAR-Inertial-Visual Odometry with Subspace-Aware Degeneracy Handling
图 1 · 摘自论文原文
  • 根据信息空间特征动态调节各传感器贡献,避免冗余优化
  • 在29个公开数据集上精度达顶尖水平,且在恶劣环境下不发散
  • 低功耗设备上每帧仅需12.3毫秒,内存占用降低6倍以上

紧耦合激光雷达-惯性-视觉里程计(LIVO)融合几何深度与视觉信息,但外部传感器会独立失效:激光雷达在扫描几何约束不足时,视觉在光照差或无纹理时。现有应对策略(二值退化检测、协方差膨胀、场景级质量门控)作用于模态层面,单一全局增益会使视觉残差进入激光雷达已约束的方向,无法聚焦于薄弱环节。本文提出子空间感知的激光雷达-惯性-视觉里程计(SA-LIVO),其子空间感知信息融合(SAIF)对联合激光雷达-视觉信息矩阵进行特征分解,通过单阈值线性截断逐方向门控,衰减低幅值方向,保留强观测方向完整强度;实现残差级鲁棒门控与场景级质量因子,剔除异常测量。激光雷达与视觉残差共享一个不变扩展卡尔曼滤波器(InEKF)回路与线性化点,使光度雅可比矩阵只需计算一次并复用于迭代。在29个公开基准序列(HILTI'22、新学院数据集(NCD)、牛津尖塔)及额外并发退化场景中,SA-LIVO精度媲美最强基线,且在其他系统发散时仍保持稳定。在所有基线均完成的HILTI'22子集上,其每帧平均耗时12.3毫秒(笔记本CPU)、26.8毫秒(嵌入式ARM板无GPU),峰值内存降低3.6至6.3倍。

原文摘要 · Abstract (English)

Tightly coupled LiDAR-inertial-visual odometry (LIVO) fuses geometric depth with visual measurements, but its exteroceptive sensors fail independently: LiDAR when scan geometry is under-constrained, vision under poor illumination or texture absence. Existing countermeasures (binary degeneracy detection, covariance inflation, scene-level quality gating) act at the modality level, so a single isotropic gain sends visual residuals into directions LiDAR already constrains well and cannot concentrate them where constraints are deficient. We propose Subspace-Aware LiDAR-inertial-visual odometry (SA-LIVO), whose Subspace-Aware Information Fusion (SAIF) eigendecomposes the joint LiDAR-visual information matrix and gates each eigendirection by a single-threshold linear clamp, attenuating low-amplitude directions while passing well-observed ones at full strength; robust per-residual gating and a scene-level quality factor screen corrupted measurements. LiDAR and visual residuals share one invariant extended Kalman filter (InEKF) loop and linearization point, letting photometric Jacobians be assembled once and reused across iterations. On 29 public-benchmark sequences (HILTI'22, Newer College Dataset (NCD), Oxford Spires), plus additional concurrent-degradation scenarios, SA-LIVO matches the strongest baselines in accuracy and stays bounded where competing systems diverge. On the HILTI'22 subset that every baseline completes, it averages 12.3 ms per frame on a laptop CPU and 26.8 ms on an embedded ARM board without GPU, at 3.6-6.3x lower peak memory.

SLAM多模态融合高效算法嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。