用可解释的视觉模型自动分析水下生物污损视频,提升检测效率与透明度。
An interpretable approach to automating the assessment of biofouling in video footage
- 基于DINOv2 ViT构建可解释的组件特征方法(ComFe)
- 比传统CNN模型性能更优,参数量减少且结果可溯源
- 适合需要透明决策过程的海洋生物污损监管场景
生物污损——指附着于水中硬质表面的生物群落——是外来海洋物种和疾病传播的重要途径。为应对这一风险,国际船舶正被要求提供其生物污损管理措施的有效证据。验证这些措施是否有效需进行水下检查,通常由潜水员或水下遥控机器人(ROV)执行,并采集与分析大量影像资料。利用计算机视觉技术实现自动化评估可显著简化该流程。本文展示如何通过可解释的组件特征(ComFe)方法结合DINOv2视觉变压器(ViT)基础模型,高效且有效地解决此问题。ComFe在性能上优于以往不可解释的卷积神经网络(CNN)方法,同时具备更少的参数量和更高的透明度——能够识别图像中哪些区域影响分类结果,以及训练数据中的哪些图像促成该判断。所有代码、数据及模型权重均已公开发布。
原文摘要 · Abstract (English)
Biofouling$\unicode{x2013}$communities of organisms that grow on hard surfaces immersed in water$\unicode{x2013}$provides a pathway for the spread of invasive marine species and diseases. To address this risk, international vessels are increasingly being obligated to provide evidence of their biofouling management practices. Verification that these activities are effective requires underwater inspections, using divers or underwater remotely operated vehicles (ROVs), and the collection and analysis of large amounts of imagery and footage. Automated assessment using computer vision techniques can significantly streamline this process, and this work shows how this challenge can be addressed efficiently and effectively using the interpretable Component Features (ComFe) approach with a DINOv2 Vision Transformer (ViT) foundation model. ComFe is able to obtain improved performance in comparison to previous non-interpretable Convolutional Neural Network (CNN) methods, with significantly fewer weights and greater transparency$\unicode{x2013}$through identifying which regions of the image contribute to the classification, and which images in the training data lead to that conclusion. All code, data and model weights are publicly released.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。