arXiv:2606.24234cs.CVcs.RO2026-06

解决船舶环境跨场景视觉定位难题,定位误差降低60%以上。

From Open Waters to Enclosed Cabins: ProteusVPR for Cross-Scene Visual Place Recognition in Maritime Perception and Cabin Inspection

论文配图:From Open Waters to Enclosed Cabins: ProteusVPR for Cross-Scene Visual Place Recognition in Maritime Perception and Cabin Inspection
图 1 · 摘自论文原文
  • 两阶段框架:先检索后精修,融合几何与视觉信息。
  • 在8000张全景图数据集上,平均定位误差降低超60%。
  • 适合船舶巡检、海洋机器人等复杂跨场景定位任务。

自主机器人在海事环境中进行视觉定位面临跨场景感知差异的挑战。机器人需在纹理稀疏、光照剧烈变化的露天甲板与结构重复、视觉模糊的封闭舱室之间切换。现有VPR方法多针对城市或室内场景,难以跨域泛化。为此,我们提出ProteusVPR,一种两阶段检索-精修框架:第一阶段使用任意标准VPR模型进行初步检索;第二阶段引入几何-视觉估计网络,融合检索图像与前两帧时序信息,结合几何描述子、局部仿射坐标系和相机方位编码,实现精准定位。为支持该任务,我们构建了XHZ数据集——一个来自实际运行船舶的8000张全景图数据集,包含多层舱室结构、甲板过渡区域,并采用严格的查询-数据库分离以进行严谨评估。在XHZ数据集上的大量实验表明,ProteusVPR在多个VPR主干网络上均显著提升定位精度,平均定位误差降低超过60%,为复杂跨场景海事环境下的精确视觉定位提供了有效且鲁棒的解决方案。

原文摘要 · Abstract (English)

Autonomous robotic inspection in maritime environments presents unique challenges for Visual Place Recognition (VPR) due to cross-scene perceptual shifts. Robots navigating ship-borne environments must transition between visually distinct domains: open decks with sparse textures and severe illumination changes, and enclosed cabins with repetitive structures and high visual ambiguity. Existing VPR methods, designed primarily for urban or indoor scenes, fail to generalize reliably across these starkly different scenarios. To address this, we propose ProteusVPR, a two-stage retrieval-refinement framework. The first stage employs any standard VPR model for initial image retrieval. The second stage introduces a geometric-visual estimation network that fuses the retrieved image with two temporally preceding frames, incorporating geometric descriptors, a local affine coordinate system, and camera azimuth encoding to achieve precise localization. To support this task, we introduce the XHZ dataset, an 8K-panoramic ship-borne dataset collected from an operational vessel, featuring multi-floor cabin structures, deck transition zones, and strict query-database separation for rigorous evaluation. Extensive experiments on the XHZ dataset demonstrate that ProteusVPR consistently improves the localization accuracy across multiple VPR backbones, reducing mean localization error by over 60\% on average and that ProteusVPR offers an effective and robust solution for precise visual localization in challenging, cross-scene maritime environments.

视觉定位船舶巡检跨场景全景数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。