arXiv:2410.19848cs.CVcs.CL2024-10被引 5

构建首个海洋哺乳动物图像分类基准,评估大模型表现

Benchmarking Large Language Models for Image Classification of Marine Mammals

  • 构建包含65类、1423张图像的海洋哺乳动物数据集,支持多层级分类
  • 多智能体系统(MAS)在分类任务中性能优于传统模型和零样本模型
  • 适合生态研究、计算机视觉与AI评估领域的研究人员参考

随着人工智能快速发展,新一代大语言模型(LLMs)在多个应用中取得突破性进展。多模态大模型也不断进步,已有多个数据集用于评估具备视觉能力的LLMs。然而,现有数据集均未聚焦海洋哺乳动物——这一对生态平衡至关重要的类群。本文构建了一个基准数据集,包含65种海洋哺乳动物的1,423张图像,每类动物按物种级、中级、群体级等不同层次进行分类。我们评估了四类方法:(1)基于神经网络嵌入的机器学习算法,(2)主流预训练神经网络,(3)零样本模型(CLIP与LLMs),(4)一种新型基于LLM的多智能体系统(MAS)。结果表明,传统模型与大模型在不同方面各具优势,而MAS进一步提升了分类性能。数据集已开源:https://github.com/yeyimilk/LLM-Vision-Marine-Animals.git。

原文摘要 · Abstract (English)

As Artificial Intelligence (AI) has developed rapidly over the past few decades, the new generation of AI, Large Language Models (LLMs) trained on massive datasets, has achieved ground-breaking performance in many applications. Further progress has been made in multimodal LLMs, with many datasets created to evaluate LLMs with vision abilities. However, none of those datasets focuses solely on marine mammals, which are indispensable for ecological equilibrium. In this work, we build a benchmark dataset with 1,423 images of 65 kinds of marine mammals, where each animal is uniquely classified into different levels of class, ranging from species-level to medium-level to group-level. Moreover, we evaluate several approaches for classifying these marine mammals: (1) machine learning (ML) algorithms using embeddings provided by neural networks, (2) influential pre-trained neural networks, (3) zero-shot models: CLIP and LLMs, and (4) a novel LLM-based multi-agent system (MAS). The results demonstrate the strengths of traditional models and LLMs in different aspects, and the MAS can further improve the classification performance. The dataset is available on GitHub: https://github.com/yeyimilk/LLM-Vision-Marine-Animals.git.

海洋哺乳动物图像分类多模态大模型基准数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。