arXiv:2604.13058cs.CLcs.LG2026-04

韩语多模态理解新基准,专测韩国文化场景下的跨模态认知能力。

KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context

论文配图:KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
图 1 · 摘自论文原文
  • 构建韩语原生多模态评测集,覆盖9学科9视觉类别及韩语特有题型。
  • 开源模型在全集上仅42.05%准确率,顶尖私有模型在难题集达52.42%。
  • 揭示韩语文化惯例与符号理解是主要瓶颈,适合多模态系统本地化研究者。

我们提出KMMMU,一个面向韩语文化与制度背景的原生韩语多模态理解评测基准。该数据集包含3,466道韩语原生考试题,覆盖9个学科和9种视觉模态类型,并设有300项韩语特有子集及627道高难度题目。与翻译或英语中心的基准不同,KMMMU聚焦受本地规范、官方标准和学科特异性视觉格式影响的信息密集型问题。实验显示,最强开源模型在全集上准确率为42.05%,最佳私有模型在难题集达到52.42%。性能在学科间差异显著,部分学科成瓶颈,韩语特有题目的差距最高达13.43%。错误分析表明,失败主因并非推理深度不足,而是对惯例到标签的映射弱、少样本符号归纳、局部知识回忆及领域特定标准理解不足。KMMMU为超越英语中心评测提供了测试平台,助力开发更可靠的专家级现实任务系统。

原文摘要 · Abstract (English)

We introduce KMMMU, a native Korean benchmark for evaluating multimodal understanding in Korean cultural and institutional settings. KMMMU contains 3,466 questions from exams natively written in Korean, covering nine disciplines and nine visual modality categories, along with a 300-item Korean-specific subset and a hard subset of 627 questions. Unlike translated or English-centric benchmarks, KMMMU targets information-dense problems shaped by local conventions, official standards, and discipline-specific visual formats. Experiments show that the strongest open-source model reaches only 42.05% accuracy on the full set, while the best proprietary model achieves 52.42% on the hard subset. Performance varies across disciplines, with some disciplines emerging as bottlenecks, and Korean-specific questions showing gaps of up to 13.43%. Error analysis suggests that these failures stem less from insufficient reasoning depth than from weak convention-to-label mapping, few-shot symbolic induction, localized knowledge recall, and domain-specific standards understanding. KMMMU provides a testbed for multimodal evaluation beyond English-centric benchmarks and for developing more reliable systems for expert real-world tasks.

多模态理解韩语评估文化适配基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。