arXiv:2608.11425cs.CV2026-08

VLMs在水下图像重建中表现远超传统物理模型。

VLMs Win a Systematic Evaluation of Underwater Image Reconstruction

论文配图:VLMs Win a Systematic Evaluation of Underwater Image Reconstruction
图 1 · 摘自论文原文
  • 构建系统化评估流程,涵盖精度、相机移动一致性及水质影响。
  • 视觉语言模型在真实水下场景中显著优于基于物理的模型。
  • 适合关注水下视觉、跨模态学习的研究者与工程师。

水下图像复原旨在恢复出仿佛无水存在的图像。以往评估缺乏系统性。本文提出一套系统化评估流程,可衡量方法的准确性、相机移动下的重建一致性以及水体参数的影响。利用该流程评估了多种现有方法,包括基于显式但近似物理模型的方案和未显式训练物理模型的视觉-语言模型(VLMs)。结果表明,VLMs整体且显著优于基于物理的模型,可能归因于强图像先验的重要性。在真实水下场景图像上的实验结果进一步验证了评估结论。

原文摘要 · Abstract (English)

Underwater image restoration consists of recovering an image which looks like there is no water present. To date, evaluation has not been systematic. This paper describes a systematic evaluation pipeline for underwater reconstruction, which can be used to assess a method for accuracy; consistency of reconstruction over camera moves; and the effect of water parameters. We use this pipeline to evaluate a range of current procedures, from models constructed using explicit but approximate physical models of scattering to Vision-Language Models (VLMs which are not currently trained with explicit physical models). Overall, VLMs wholly and significantly outperform physically based models in our evaluation, likely because of the importance of a strong image prior. Results on images of real underwater scenes strongly confirm the evaluation.

水下图像视觉语言模型图像复原

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。