arXiv:2511.22860cs.ROcs.CV2025-11

MARVO通过物理建模与强化学习,提升水下视觉里程计的精度与鲁棒性。

MARVO: Marine-Adaptive Radiance-aware Visual Odometry

  • 融合水下成像物理模型与可微匹配,补偿颜色衰减和对比度损失。
  • 结合惯性、压力与视觉信息,在实时框架下实现全状态最大后验估计。
  • 采用强化学习优化全局轨迹,突破传统最小二乘法的局部最优陷阱。

水下视觉定位因波长相关衰减、纹理贫乏及非高斯传感器噪声而极具挑战。本文提出MARVO,一种融合物理感知与学习机制的里程计框架,整合水下图像形成建模、可微匹配与强化学习优化。前端通过引入物理感知辐射率适配器扩展基于Transformer的特征匹配器,有效补偿浊度下的颜色通道衰减与对比度损失,生成几何一致的特征对应。半稠密匹配结果与惯性及压力数据在因子图后端融合,利用GTSAM库构建关键帧驱动的视觉-惯性-气压估计算法。每个关键帧引入:(i) 预积分IMU运动因子,(ii) MARVO推导的视觉姿态因子,(iii) 气压深度先验,实现实时全状态最大后验估计。最后,设计基于强化学习的姿态图优化器,通过学习SE(2)空间中的最优回缩动作,将全局轨迹优化至超越经典最小二乘解的局部极小值。

原文摘要 · Abstract (English)

Underwater visual localization remains challenging due to wavelength-dependent attenuation, poor texture, and non-Gaussian sensor noise. We introduce MARVO, a physics-aware, learning-integrated odometry framework that fuses underwater image formation modeling, differentiable matching, and reinforcement-learning optimization. At the front-end, we extend transformer-based feature matcher with a Physics Aware Radiance Adapter that compensates for color channel attenuation and contrast loss, yielding geometrically consistent feature correspondences under turbidity. These semi dense matches are combined with inertial and pressure measurements inside a factor-graph backend, where we formulate a keyframe-based visual-inertial-barometric estimator using GTSAM library. Each keyframe introduces (i) Pre-integrated IMU motion factors, (ii) MARVO-derived visual pose factors, and (iii) barometric depth priors, giving a full-state MAP estimate in real time. Lastly, we introduce a Reinforcement-Learningbased Pose-Graph Optimizer that refines global trajectories beyond local minima of classical least-squares solvers by learning optimal retraction actions on SE(2).

水下定位视觉里程计强化学习物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。