arXiv:2507.02148cs.CV2025-07被引 1

用合成数据微调模型,提升水下单目测距精度

Underwater Monocular Metric Depth Estimation: Real-World Benchmarks and Synthetic Fine-Tuning with Vision Foundation Models

  • 用物理模型生成合成水下图像,微调深度模型
  • 微调后模型在多个真实水下数据集上性能显著提升
  • 适合做水下视觉感知与机器人导航的研究者

单目深度估计已从序数深度进展到度量深度,但在水下环境中仍受限于光衰减、散射、色彩失真、浑浊度以及高质量度量真值数据的缺乏。本文构建了首个基于真实水下数据集(FLSea 和 SQUID)的零样本与微调模型综合评测基准,评估多种前沿视觉基础模型在不同水下条件和深度范围下的表现。结果显示,仅在陆地数据(真实或合成)上训练的大模型在空中场景有效,但水下域偏移严重导致性能下降。为此,我们采用基于物理的水下图像形成模型,生成合成水下变体的 Hypersim 数据集,以 ViT-S 为编码器对 Depth Anything V2 进行微调。微调后模型在所有基准上均一致提升,优于仅在干净陆地 Hypersim 上训练的基线。本研究揭示了领域适应与尺度感知监督对基础模型在复杂环境下实现鲁棒、泛化度量深度预测的重要性。

原文摘要 · Abstract (English)

Monocular depth estimation has recently progressed beyond ordinal depth to provide metric depth predictions. However, its reliability in underwater environments remains limited due to light attenuation and scattering, color distortion, turbidity, and the lack of high-quality metric ground truth data. In this paper, we present a comprehensive benchmark of zero-shot and fine-tuned monocular metric depth estimation models on real-world underwater datasets with metric depth annotations, including FLSea and SQUID. We evaluated a diverse set of state-of-the-art Vision Foundation Models across a range of underwater conditions and depth ranges. Our results show that large-scale models trained on terrestrial data (real or synthetic) are effective in in-air settings, but perform poorly underwater due to significant domain shifts. To address this, we fine-tune Depth Anything V2 with a ViT-S backbone encoder on a synthetic underwater variant of the Hypersim dataset, which we simulated using a physically based underwater image formation model. Our fine-tuned model consistently improves performance across all benchmarks and outperforms baselines trained only on the clean in-air Hypersim dataset. This study presents a detailed evaluation and visualization of monocular metric depth estimation in underwater scenes, emphasizing the importance of domain adaptation and scale-aware supervision for achieving robust and generalizable metric depth predictions using foundation models in challenging environments.

水下视觉深度估计域适应基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。