arXiv:2410.21259cs.CVcs.AI2024-10被引 22

用大模型自动生成测评数据,自动评估视觉语言模型性能。

AutoBench-V: Can Large Vision-Language Models Benchmark Themselves?

  • 用文本生成图像,再让模型自问自答完成测评。
  • 在5种能力上评测9个主流模型,结果可靠有效。
  • 适合需要快速评估模型性能的研究者使用。

大型视觉语言模型(LVLMs)在融合视觉与语言信息方面日益重要。然而,现有评估基准构建成本高、依赖人工,且一旦建立便难以调整。尽管文本模态已有自动评估探索,但视觉模态仍鲜有研究。本文提出:能否让大模型自我评估?为此,我们设计了AutoBench-V——一个按需生成的自动化评估框架,通过文本到图像模型生成相关图像样本,并利用LVLM协同完成视觉问答任务,实现高效灵活的评估。在五类用户指定能力下对九个主流LVLM进行评估,结果表明该框架具有有效性与可靠性。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) have become essential for advancing the integration of visual and linguistic information. However, the evaluation of LVLMs presents significant challenges as the evaluation benchmark always demands lots of human cost for its construction, and remains static, lacking flexibility once constructed. Even though automatic evaluation has been explored in textual modality, the visual modality remains under-explored. As a result, in this work, we address a question: "Can LVLMs themselves be used to benchmark each other in the visual automatically domain?". We introduce AutoBench-V, an automated framework for serving evaluation on demand, i.e., benchmarking LVLMs based on specific aspects of model capability. AutoBench-V leverages text-to-image models to generate relevant image samples and then utilizes LVLMs to orchestrate visual question-answering (VQA) tasks, completing the evaluation process efficiently and flexibly. Through an extensive evaluation of nine popular LVLMs across five demanded user inputs (i.e., evaluation capabilities), the framework shows effectiveness and reliability.

视觉语言模型自动评估自举测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。