arXiv:2412.15574cs.CV2024-12被引 1

构建首个深海生物多模态大模型评测基准,测试模型对深海物种的理解能力。

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

  • 基于真实深海图像设计问答任务,评估多模态LLM对深海生物的识别与理解。
  • 100张图像配日本语问答,顶尖模型o1仅达50%准确率,说明当前水平仍不足。
  • 适合海洋科研、多模态模型开发者关注,推动深海智能认知发展。

日本海洋地球科学与技术机构(JAMSTEC)发布了深海影像数据集J-EDI(https://www.godac.jamstec.go.jp/jedi/e/index.html),包含大量深海生物、海底地貌及物理过程的视频与图像。本文提出J-EDI QA,一个用于评估多模态大语言模型(LLM)理解深海生物图像能力的基准。该基准包含100张图像,每张配有由JAMSTEC研究人员设计的日本语问答对,含四个选项。在本文评估中,OpenAI o1模型在该任务上仅获得50%的正确率,表明即使截至2024年12月,现有顶级模型对深海物种的理解仍未达到专家水平,亟需发展更专精的深海生物多模态大模型。

原文摘要 · Abstract (English)

Japan Agency for Marine-Earth Science and Technology (JAMSTEC) has made available the JAMSTEC Earth Deep-sea Image (J-EDI), a deep-sea video and image archive (https://www.godac.jamstec.go.jp/jedi/e/index.html). This archive serves as a valuable resource for researchers and scholars interested in deep-sea imagery. The dataset comprises images and videos of deep-sea phenomena, predominantly of marine organisms, but also of the seafloor and physical processes. In this study, we propose J-EDI QA, a benchmark for understanding images of deep-sea organisms using a multimodal large language model (LLM). The benchmark is comprised of 100 images, accompanied by questions and answers with four options by JAMSTEC researchers for each image. The QA pairs are provided in Japanese, and the benchmark assesses the ability to understand deep-sea species in Japanese. In the evaluation presented in this paper, OpenAI o1 achieved a 50% correct response rate. This result indicates that even with the capabilities of state-of-the-art models as of December 2024, deep-sea species comprehension is not yet at an expert level. Further advances in deep-sea species-specific LLMs are therefore required.

多模态模型深海生物评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。