arXiv:2604.00909cs.CV2026-04被引 2

针对日语视觉问答数据集质量差的问题,构建高质量评测集JAMMEval。

JAMMEval: A Refined Collection of Japanese Benchmarks for Reliable VLM Evaluation

  • 通过两轮人工标注系统性优化7个日语文本-图像基准数据集
  • 新评测集评分更准确、波动小,能更好区分模型能力差异
  • 适合研究日语多模态模型或追求可靠评测的团队使用

可靠的评估对视觉语言模型(VLM)的发展至关重要。然而,与英语基准相比,日语视觉问答(VQA)基准尚未经历充分的迭代优化。现有许多基准存在问题,如问题模糊、答案错误,以及无需视觉理解即可解题的样本,严重损害了评估可靠性,导致模型比较产生误导性结论。为此,我们提出JAMMEval,一个经过系统性优化的日语基准集合,通过两轮人工标注对七个现有日语文本-图像基准数据集进行改进,显著提升数据质量与评估可靠性。实验中,我们在JAMMEval上评估了开源与专有VLM,并分析近期模型在日语VQA上的表现。结果表明,经优化后的基准能更真实反映模型能力,减少运行间方差,并增强对不同能力水平模型的区分力。我们已公开数据集与代码,以推动日语VLM的可靠评估。

原文摘要 · Abstract (English)

Reliable evaluation is essential for the development of vision-language models (VLMs). However, Japanese VQA benchmarks have undergone far less iterative refinement than their English counterparts. As a result, many existing benchmarks contain issues such as ambiguous questions, incorrect answers, and instances that can be solved without visual grounding, undermining evaluation reliability and leading to misleading conclusions in model comparisons. To address these limitations, we introduce JAMMEval, a refined collection of Japanese benchmarks for reliable VLM evaluation. It is constructed by systematically refining seven existing Japanese benchmark datasets through two rounds of human annotation, improving both data quality and evaluation reliability. In our experiments, we evaluate open-weight and proprietary VLMs on JAMMEval and analyze the capabilities of recent models on Japanese VQA. We further demonstrate the effectiveness of our refinement by showing that the resulting benchmarks yield evaluation scores that better reflect model capability, exhibit lower run-to-run variance, and improve the ability to distinguish between models of different capability levels. We release our dataset and code to advance reliable evaluation of VLMs.

视觉问答日语多模态评测集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。