arXiv:2608.19710cs.CVcs.AI2026-08

用声呐增强视觉模型,让水下机器人在恶劣环境下仍能可靠感知。

Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions

论文配图:Robust Cross-Modal Foundation Model Perception for Underwater Robots under Degraded Visual Conditions
图 1 · 摘自论文原文
  • 冻结预训练视觉模型,动态融合声呐信息适应不同退化条件。
  • 极端退化下检测准确率从46.1%提升至61.5%,相对提高33.5%。
  • 适合水下机器人、海洋探测等复杂环境下的多模态感知研究。

可靠的水下机器人感知仍具挑战,因浊度、波长衰减、光照不足、散射和模糊导致光学图像严重退化。尽管声呐提供受光学影响较小的互补信息,但以往研究多聚焦特征对齐与正常条件下的检测性能。本文研究视觉可靠性下降时的跨模态鲁棒性,评估预训练视觉基础模型表示能否通过声呐补充。采用冻结的DINOv2作为视觉编码器,构建从清晰到极端退化的五级可控基准测试。对比传统视觉检测、冻结基础模型表示、声呐上下文、固定多模态融合、清洁数据训练的自适应门控以及退化感知门控融合。所提方法在全退化范围内训练融合机制,保持视觉与声呐编码器冻结,使模态贡献自适应调整而无需微调预训练主干。在极端综合退化下,DINOv2基线平衡准确率为0.4610,退化感知视觉-声呐融合达到0.6152,相对提升33.5%。学习到的声呐贡献从清洁条件下14.2%增至极端退化下41.3%,表明跨模态依赖可自适应重分配。融合在严重浊度与模糊时增益最大,而仅颜色衰减带来有限收益。结果表明,基础模型表示在严重信息丢失下仍有价值但不足,显式适应融合以模态可靠性可显著提升水下多模态感知鲁棒性。

原文摘要 · Abstract (English)

Reliable underwater robotic perception remains difficult because optical imagery degrades under turbidity, wavelength-dependent attenuation, low illumination, scattering, and blur. Although sonar provides complementary information that is less affected by optical visibility, prior visual-sonar research has largely focused on feature alignment and nominal detection performance. We investigate cross-modal robustness as visual reliability deteriorates and assess whether pretrained visual foundation-model representations can be complemented by sonar under severe degradation. We use frozen DINOv2 as the visual encoder and construct a controlled five-level benchmark ranging from clean to extreme visual conditions. We compare conventional visual detection, frozen foundation-model representations, sonar context, fixed multimodal fusion, clean-trained adaptive gating, and degradation-aware gated fusion. Our method trains the fusion mechanism across the full range of degradation while keeping the visual and sonar encoders frozen, allowing modality contributions to adapt without fine-tuning the pretrained backbone. Under extreme combined degradation, the DINOv2 baseline achieves 0.4610 balanced accuracy, while degradation-aware visual-sonar fusion reaches 0.6152, a 33.5% relative improvement. The learned sonar contribution increases from 14.2% under clean conditions to 41.3% under extreme degradation, demonstrating adaptive redistribution of cross-modal reliance. Fusion provides the largest gains under severe turbidity and blur, whereas color attenuation alone yields little additional benefit. These results show that foundation-model representations remain valuable but insufficient under severe information loss, and that explicitly adapting fusion to modality reliability can improve robust underwater multimodal perception.

水下感知多模态融合基础模型声呐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。