首个面向越南语的多任务多模态理解基准,测试模型跨模态推理能力。
VMMU: A Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark
- 构建2.5千个跨7类任务的越南语多模态问答数据集
- 主流模型在该基准上平均准确率仅66%,低于预期
- 核心瓶颈是跨模态对齐与推理,而非文字识别能力
我们提出VMMU,一个面向越南语的多任务多模态理解与推理基准,用于评估视觉语言模型(VLMs)在非英语场景下对视觉与文本信息的解读和推理能力。VMMU包含2.5k个跨7个任务的多模态问题,涵盖科学数学解题、数据解读、规则驱动的视觉推理及抽象视觉推理等多种情境。所有题目均要求真实多模态融合,杜绝仅依赖文本或基于OCR的捷径。我们在一系列先进的专有及开源VLMs上评估了该基准。尽管越南语光学字符识别(OCR)表现良好,但专有模型的平均准确率仅为66%。进一步分析表明,失败主因并非OCR,而是文本与视觉证据之间的多模态定位与推理能力不足。代码与数据已公开于https://vmmu-bench.github.io/
原文摘要 · Abstract (English)
We introduce VMMU, a Vietnamese Multitask Multimodal Understanding and Reasoning Benchmark designed to evaluate how vision-language models (VLMs) interpret and reason over visual and textual information beyond English. VMMU consists of 2.5k multimodal questions across 7 tasks, covering a diverse range of problem contexts, including STEM problem solving, data interpretation, rule-governed visual reasoning, and abstract visual reasoning. All questions require genuine multimodal integration, rather than reliance on text-only cues or OCR-based shortcuts. We evaluate a diverse set of state-of-the-art proprietary and open-source VLMs on VMMU. Despite strong Vietnamese OCR performance, proprietary models achieve only 66% mean accuracy. Further analysis shows that the primary source of failure is not OCR, but instead multimodal grounding and reasoning over text and visual evidence. Code and data are available at https://vmmu-bench.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。