首个海洋领域多维度视觉语言模型评测基准,揭示现有模型在专业海洋问题上表现不足。
MarineEval: Assessing the Marine Intelligence of Vision-Language Models
- 构建首个覆盖7类任务、20种能力的海洋视觉问答数据集,含2000组图文对。
- 17个主流VLM在该数据集上表现不佳,平均准确率未达人类水平。
- 专为海洋领域设计,适合研究海洋智能与跨模态模型泛化能力的学者。
尽管大语言模型(LLMs)和视觉语言模型(VLMs)在通用任务中取得显著进展,但它们是否能作为海洋领域的专家准确回答专业问题仍存疑问。本文提出首个大规模海洋视觉语言模型评测基准MarineEval,包含2000个基于图像的问答对,涵盖7个任务维度和20个能力维度,所有数据均经海洋领域专家审核验证。我们对17个现有VLM进行了全面评估,结果表明当前模型在应对海洋特定问题时表现有限,存在显著提升空间。本工作旨在推动海洋智能与跨模态理解的研究发展。
原文摘要 · Abstract (English)
We have witnessed promising progress led by large language models (LLMs) and further vision language models (VLMs) in handling various queries as a general-purpose assistant. VLMs, as a bridge to connect the visual world and language corpus, receive both visual content and various text-only user instructions to generate corresponding responses. Though great success has been achieved by VLMs in various fields, in this work, we ask whether the existing VLMs can act as domain experts, accurately answering marine questions, which require significant domain expertise and address special domain challenges/requirements. To comprehensively evaluate the effectiveness and explore the boundary of existing VLMs, we construct the first large-scale marine VLM dataset and benchmark called MarineEval, with 2,000 image-based question-answering pairs. During our dataset construction, we ensure the diversity and coverage of the constructed data: 7 task dimensions and 20 capacity dimensions. The domain requirements are specially integrated into the data construction and further verified by the corresponding marine domain experts. We comprehensively benchmark 17 existing VLMs on our MarineEval and also investigate the limitations of existing models in answering marine research questions. The experimental results reveal that existing VLMs cannot effectively answer the domain-specific questions, and there is still a large room for further performance improvements. We hope our new benchmark and observations will facilitate future research. Project Page: http://marineeval.hkustvgd.com/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。