用视觉模型和大模型自动评估船体生物污损等级,提升检测效率与可解释性。
Automated Marine Biofouling Assessment: Benchmarking Computer Vision and Multimodal LLMs on the Level of Fouling Scale
- 结合卷积网络与多模态大模型,通过结构化提示实现零样本分类。
- 大模型在未训练情况下达到媲美视觉模型的准确率,尤其擅长中间等级判断。
- 适合海洋生态监测、船舶维护等需要高效可解释评估的场景。
船体生物污损带来重大生态、经济与生物安全风险。传统人工潜水检测危险且难以规模化。本研究利用定制计算机视觉模型与大规模多模态语言模型(LLMs),在新西兰农业部提供的专家标注数据集上,对生物污损严重程度的等级(LoF)进行自动化分类。评估了卷积神经网络、基于Transformer的分割模型及零样本大模型。视觉模型在极端污损等级上表现优异,但在中间等级因数据分布不均与图像构图问题表现不佳。大模型在无训练条件下通过结构化提示与检索机制实现竞争力性能,并提供可解释输出。结果表明两类方法互补,融合分割覆盖与大模型推理的混合方案是实现可扩展、可解释生物污损评估的可行路径。
原文摘要 · Abstract (English)
Marine biofouling on vessel hulls poses major ecological, economic, and biosecurity risks. Traditional survey methods rely on diver inspections, which are hazardous and limited in scalability. This work investigates automated classification of biofouling severity on the Level of Fouling (LoF) scale using both custom computer vision models and large multimodal language models (LLMs). Convolutional neural networks, transformer-based segmentation, and zero-shot LLMs were evaluated on an expert-labelled dataset from the New Zealand Ministry for Primary Industries. Computer vision models showed high accuracy at extreme LoF categories but struggled with intermediate levels due to dataset imbalance and image framing. LLMs, guided by structured prompts and retrieval, achieved competitive performance without training and provided interpretable outputs. The results demonstrate complementary strengths across approaches and suggest that hybrid methods integrating segmentation coverage with LLM reasoning offer a promising pathway toward scalable and interpretable biofouling assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。