arXiv:2503.23730cs.CVcs.AI2025-03CVPR被引 5

首个面向韩语大视觉语言模型的开放式问答基准,实现客观可靠评估。

KOFFVQA: An Objectively Evaluated Free-form VQA Benchmark for Large Vision-Language Models in the Korean Language

  • 构建10维度评分标准,用规则化流程替代主观打分。
  • 275个精心设计的韩语图文问答对,支持自由回答形式。
  • 开源评估代码,小模型也可可靠评测大模型表现。

大型视觉语言模型(VLMs)的兴起催生了多种评估基准,但现有方法普遍存在两大问题:要么要求模型从预设答案中选择,牺牲开放性;要么依赖评判模型打分,导致结果主观不可靠。此外,韩语领域缺乏专用评估基准,而语言差异会影响生成模型性能。为此,我们提出KOFFVQA——一个面向韩语大视觉语言模型的通用开放式图文问答基准。该基准包含275个精心设计的问题,每个问题配有一张图像和涵盖10个维度的评分标准。通过预定义的规则化评分机制,可避免主观偏差,使小型开源模型也能可靠评估大型模型。我们在基准上测试了大量现有VLM,并实验证明,基于预设评分标准的方法比传统方法更可靠。评估代码已开源:https://github.com/maum-ai/KOFFVQA。

原文摘要 · Abstract (English)

The recent emergence of Large Vision-Language Models(VLMs) has resulted in a variety of different benchmarks for evaluating such models. Despite this, we observe that most existing evaluation methods suffer from the fact that they either require the model to choose from pre-determined responses, sacrificing open-endedness, or evaluate responses using a judge model, resulting in subjective and unreliable evaluation. In addition, we observe a lack of benchmarks for VLMs in the Korean language, which are necessary as a separate metric from more common English language benchmarks, as the performance of generative language models can differ significantly based on the language being used. Therefore, we present KOFFVQA, a general-purpose free-form visual question answering benchmark in the Korean language for the evaluation of VLMs. Our benchmark consists of 275 carefully crafted questions each paired with an image and grading criteria covering 10 different aspects of VLM performance. The grading criteria eliminate the problem of unreliability by allowing the judge model to grade each response based on a pre-determined set of rules. By defining the evaluation criteria in an objective manner, even a small open-source model can be used to evaluate models on our benchmark reliably. In addition to evaluating a large number of existing VLMs on our benchmark, we also experimentally verify that our method of using pre-existing grading criteria for evaluation is much more reliable than existing methods. Our evaluation code is available at https://github.com/maum-ai/KOFFVQA

视觉问答韩语基准测试客观评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。