利用位置信息提升海底图像自监督学习效果,显著改善分类性能。
Investigating Location-Regularised Self-Supervised Feature Learning for Seafloor Visual Imagery
- 引入位置正则化增强自监督学习,结合地理坐标优化特征提取。
- 在三种数据集上平均提升F1分数4.9%(CNN)和6.3%(ViT)。
- 低维表示下位置正则化效果更明显,适合资源受限场景使用。
机器人采集的海底视觉图像高效解析可提升海洋监测与探索效率。尽管已有研究指出位置元数据能增强自监督学习(SSL),但其在不同SSL策略、模型及数据集上的效果仍待深入。本研究评估了六种先进SSL框架中位置正则化的性能,涵盖具有不同隐空间维度的卷积神经网络(CNN)与视觉变换器(ViT)。在三个多样化的海底图像数据集上,位置正则化均优于标准SSL,CNN平均F1得分提升$4.9 \pm 4.0\%$,ViT提升$6.3 \pm 8.9\%$。CNN在通用数据集预训练时受益于高维隐空间,而数据集优化的SSL在512与128维间表现相当。位置正则化使CNN性能相比预训练模型分别提升$2.7 \pm 2.7\%$(高维)与$10.1 \pm 9.4\%$(低维)。对ViT而言,高维表示对预训练与数据集优化均有利。尽管位置正则化优于标准方法,预训练ViT已展现强泛化能力,与最优位置正则化方法的F1分数分别为$0.795 \pm 0.075$与$0.795 \pm 0.077$。结果表明位置信息对SSL正则化价值显著,尤其在低维表示中;同时高维ViT在海底图像分析中具备优异泛化性。
原文摘要 · Abstract (English)
High-throughput interpretation of robotically gathered seafloor visual imagery can increase the efficiency of marine monitoring and exploration. Although recent research has suggested that location metadata can enhance self-supervised feature learning (SSL), its benefits across different SSL strategies, models and seafloor image datasets are underexplored. This study evaluates the impact of location-based regularisation on six state-of-the-art SSL frameworks, which include Convolutional Neural Network (CNN) and Vision Transformer (ViT) models with varying latent-space dimensionality. Evaluation across three diverse seafloor image datasets finds that location-regularisation consistently improves downstream classification performance over standard SSL, with average F1-score gains of $4.9 \pm 4.0%$ for CNNs and $6.3 \pm 8.9%$ for ViTs, respectively. While CNNs pretrained on generic datasets benefit from high-dimensional latent representations, dataset-optimised SSL achieves similar performance across the high (512) and low (128) dimensional latent representations. Location-regularised SSL improves CNN performance over pre-trained models by $2.7 \pm 2.7%$ and $10.1 \pm 9.4%$ for high and low-dimensional latent representations, respectively. For ViTs, high-dimensionality benefits both pre-trained and dataset-optimised SSL. Although location-regularisation improves SSL performance compared to standard SSL methods, pre-trained ViTs show strong generalisation, matching the best-performing location-regularised SSL with F1-scores of $0.795 \pm 0.075$ and $0.795 \pm 0.077$, respectively. The findings highlight the value of location metadata for SSL regularisation, particularly when using low-dimensional latent representations, and demonstrate strong generalisation of high-dimensional ViTs for seafloor image analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。