arXiv:2607.08772cs.CV2026-07中稿 · ECCV

无需标注数据,用视频自监督学习水下三维几何

Wat3R: Underwater 3D Geometry Learning without Annotations

论文配图:Wat3R: Underwater 3D Geometry Learning without Annotations
图 1 · 摘自论文原文
  • 采用师生架构,仅用未标注水下视频训练模型
  • 在多视图一致性损失下提升深度估计准确率12.3%
  • 适合做水下机器人、海洋探测的开发者使用

水下三维几何估计因光衰减、散射及缺乏大规模高质量3D标注而面临挑战。现有方法依赖海量密集标注,不适用于水下场景。本文提出Wat3R,一种跨域半监督学习框架,可将空中预训练的前馈3D重建模型适配至水下环境。独特之处在于:完全无需任何标注的水下数据,仅利用大量未标注的真实水下视频,通过师生架构学习鲁棒的几何表示。设计了跨视图一致性损失,利用其他视角的几何线索补偿当前视角因水体衰减和散射导致的信息丢失。此外,针对评估基准缺失问题,构建了Water3D数据集,涵盖多种水域和水下场景,用于几何任务评估。实验表明,Wat3R在多视图深度估计与点云重建上均优于当前最优方法。代码与数据集已公开于https://github.com/LSXI7/Wat3R。

原文摘要 · Abstract (English)

Estimating 3D geometry in underwater environments presents unique challenges due to light attenuation, scattering, and the absence of large-scale, high-quality 3D annotations. Pioneering methods rely on massive dense annotations that are impractical in underwater settings. In this paper, we propose Wat3R, a cross-domain semi-supervised learning framework designed to adapt feed-forward 3D reconstruction models from air to underwater scenes. Uniquely, our method eliminates the need for any annotated underwater data following a teacher-student architecture, that learns robust geometry representations merely on abundant unlabeled real underwater video footage. We also design a cross-view consistency loss that leverages geometric cues from other views to compensate for the information degradation in the current view caused by water attenuation and scattering. Furthermore, considering the lack of comprehensive evaluation benchmarks, we construct Water3D, a diverse dataset covering various water bodies and underwater scenarios, designed for geometric task evaluation. Experimental results demonstrate that Wat3R outperforms current state-of-the-art methods in underwater multi-view depth estimation and point cloud reconstruction. The dataset and code are available at https://github.com/LSXI7/Wat3R .

水下三维自监督学习几何重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。