arXiv:2512.14620cs.CLcs.AI2025-12被引 1

构建日语多学科图文理解基准,提升模型视觉文本融合能力评估标准。

JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction

  • 用生成模型产图题,人工校验优化,低成本构建高质量图文数据集。
  • 开源大模型在新基准上表现差,凸显其日语图文理解短板。
  • 适合评估日语多模态模型,为未来图像问答基准建设提供方法参考。

本文提出JMMMU-Pro,一个基于图像的日本多学科多模态理解基准,以及Vibe基准构建法,一种可扩展的构建方法。继MMMU到MMMU-Pro的发展,JMMMU-Pro通过将问题图像与文本合并为单一图像,要求模型通过视觉感知实现跨模态理解。构建过程中,采用图像生成模型(如Nano Banana Pro)生成候选视觉问题,人类验证并按需调整提示词重生成以保证质量。借助Nano Banana Pro的高真实感生成能力和清晰日文嵌入能力,实现了低成本、多背景布局的高质量基准构建。实验表明,所有开源大模型在JMMMU-Pro上表现显著不足,凸显其作为开放社区重要评估基准的价值。我们认为JMMMU-Pro提供了更严格的日语多模态模型评估工具,而Vibe基准构建法也为未来图像问答基准开发提供了高效指导。

原文摘要 · Abstract (English)

This paper introduces JMMMU-Pro, an image-based Japanese Multi-discipline Multimodal Understanding Benchmark, and Vibe Benchmark Construction, a scalable construction method. Following the evolution from MMMU to MMMU-Pro, JMMMU-Pro extends JMMMU by composing the question image and question text into a single image, thereby creating a benchmark that requires integrated visual-textual understanding through visual perception. To build JMMMU-Pro, we propose Vibe Benchmark Construction, a methodology in which an image generative model (e.g., Nano Banana Pro) produces candidate visual questions, and humans verify the outputs and, when necessary, regenerate with adjusted prompts to ensure quality. By leveraging Nano Banana Pro's highly realistic image generation capabilities and its ability to embed clean Japanese text, we construct a high-quality benchmark at low cost, covering a wide range of background and layout designs. Experimental results show that all open-source LMMs struggle substantially with JMMMU-Pro, underscoring JMMMU-Pro as an important benchmark for guiding future efforts in the open-source community. We believe that JMMMU-Pro provides a more rigorous evaluation tool for assessing the Japanese capabilities of LMMs and that our Vibe Benchmark Construction also offers an efficient guideline for future development of image-based VQA benchmarks.

多模态日语理解图像问答基准构建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。