首个面向日语文化的多模态理解基准,评测大模型对日本文化的理解能力。
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation
- 构建双子集评测框架:去文化与文化特异性任务并行
- 发现多数模型在日语中表现下降,且文化理解严重不足
- 适合关注日语多模态模型与跨文化评估的研究者
非英语语言的大规模多模态模型(LMMs)研究亟需加速,以提升更广泛人群的用户体验。本文提出JMMMU(日本MMMUs),首个基于日本文化背景的大型日语多模态理解评测基准。该基准包含两个互补子集:(i) 去文化(CA)子集,选取数学等文化无关学科并翻译为日语,可与英文版MMMU一对一比较;(ii) 文化特异性(CS)子集,全新设计反映日本文化的内容。通过CA子集发现,许多模型在日语评测中性能下降,仅由语言差异导致;通过CS子集揭示其日本文化理解能力不足。结合两者发现,部分模型在CA子集表现良好但在CS子集表现差,暴露了对日语浅层理解、缺乏文化深度的问题。本工作旨在推动日语下LMM发展,并为多语言多文化基准建设提供范例。项目主页:https://mmmu-japanese-benchmark.github.io/JMMMU/
原文摘要 · Abstract (English)
Accelerating research on Large Multimodal Models (LMMs) in non-English languages is crucial for enhancing user experiences across broader populations. In this paper, we introduce JMMMU (Japanese MMMU), the first large-scale Japanese benchmark designed to evaluate LMMs on expert-level tasks based on the Japanese cultural context. To facilitate comprehensive culture-aware evaluation, JMMMU features two complementary subsets: (i) culture-agnostic (CA) subset, where the culture-independent subjects (e.g., Math) are selected and translated into Japanese, enabling one-to-one comparison with its English counterpart MMMU; and (ii) culture-specific (CS) subset, comprising newly crafted subjects that reflect Japanese cultural context. Using the CA subset, we observe performance drop in many LMMs when evaluated in Japanese, which is purely attributable to language variation. Using the CS subset, we reveal their inadequate Japanese cultural understanding. Further, by combining both subsets, we identify that some LMMs perform well on the CA subset but not on the CS subset, exposing a shallow understanding of the Japanese language that lacks depth in cultural understanding. We hope this work will not only help advance LMM performance in Japanese but also serve as a guideline to create high-standard, culturally diverse benchmarks for multilingual LMM development. The project page is https://mmmu-japanese-benchmark.github.io/JMMMU/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。