arXiv:2409.09269cs.CVcs.AI2024-09中稿 · The First Workshop…被引 19

为跨任务选对视觉语言模型提供系统方法

Guiding Vision-Language Model Selection for Visual Question-Answering Across Tasks, Domains, and Knowledge Types

  • 构建VQA360数据集,覆盖多种任务类型与知识类型
  • 提出GoEval评估指标,与人工判断相关性达56.71%
  • 揭示无通用最优模型,助开发者精准选型

视觉问答(VQA)已成为提升用户体验的关键,尤其在视觉语言模型(VLMs)具备更强泛化能力后。但在实际应用中,如何通过标准化框架评估VLM以满足特定需求仍具挑战。本文提出端到端解决方案:构建VQA360——一个源自主流VQA基准的新型数据集,包含任务类型、应用领域和知识类型标注,支持全面评估。同时引入基于GPT-4o开发的多模态评估指标GoEval,与人工判断的相关性达56.71%。对主流VLM的实验表明,不存在通用最优模型,模型选择成为关键设计决策。专有模型如Gemini-1.5-Pro和GPT-4o-mini整体表现更优,但开源模型InternVL-2-8B和CogVLM-2-Llama-3-19B也展现出竞争力,且具额外优势。该框架可拓展至其他任务。

原文摘要 · Abstract (English)

Visual Question-Answering (VQA) has become key to user experience, particularly after improved generalization capabilities of Vision-Language Models (VLMs). But evaluating VLMs for an application requirement using a standardized framework in practical settings is still challenging. This paper aims to solve that using an end-to-end framework. We present VQA360 - a novel dataset derived from established VQA benchmarks, annotated with task types, application domains, and knowledge types, for a comprehensive evaluation. We also introduce GoEval, a multimodal evaluation metric developed using GPT-4o, achieving a correlation factor of 56.71% with human judgments. Our experiments with state-of-the-art VLMs reveal that no single model excels universally, thus, making a right choice a key design decision. Proprietary models such as Gemini-1.5-Pro and GPT-4o-mini generally outperform others, but open-source models like InternVL-2-8B and CogVLM-2-Llama-3-19B also demonstrate competitive strengths, while providing additional advantages. Our framework can also be extended to other tasks.

视觉问答模型评估多模态VLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。