评测大模型在鱼类识别中的表现,发现准确率不足10%。
FishNet++: Analyzing the capabilities of Multimodal Large Language Models in marine biology
- 构建鱼类多模态数据集FishNet++,含超35万文本描述
- 最佳开源模型对鱼类物种识别准确率低于10%
- 适合海洋生态监测与专用视觉语言模型研究者
多模态大语言模型在跨领域任务中表现出色,但在海洋生物学等专业领域的应用仍缺乏深入研究。本文系统评估了当前先进多模态大模型,发现其在细粒度鱼类物种识别上存在显著局限,最佳开源模型准确率不足10%,而该任务对受人为压力影响的海洋生态系统监测至关重要。为填补这一空白并探究性能瓶颈是否源于领域知识缺失,我们提出了FishNet++——一个大规模多模态基准。该数据集包含35,133条文本描述、706,426个关键点标注和119,399个边界框,显著扩展了现有资源。通过提供全面标注,本工作推动了面向水生科学的专用视觉-语言模型的发展与评估。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have demonstrated impressive cross-domain capabilities, yet their proficiency in specialized scientific fields like marine biology remains underexplored. In this work, we systematically evaluate state-of-the-art MLLMs and reveal significant limitations in their ability to perform fine-grained recognition of fish species, with the best open-source models achieving less than 10\% accuracy. This task is critical for monitoring marine ecosystems under anthropogenic pressure. To address this gap and investigate whether these failures stem from a lack of domain knowledge, we introduce FishNet++, a large-scale, multimodal benchmark. FishNet++ significantly extends existing resources with 35,133 textual descriptions for multimodal learning, 706,426 key-point annotations for morphological studies, and 119,399 bounding boxes for detection. By providing this comprehensive suite of annotations, our work facilitates the development and evaluation of specialized vision-language models capable of advancing aquatic science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。