构建多语言多模态公务员考试数据集,挑战视觉语言模型真实场景推理能力。
EuraGovExam: A Multilingual Multimodal Benchmark from Real-World Civil Service Exams
- 将题目、选项与图像整合为单张图片,要求模型直接从视觉输入中理解布局与语义。
- 包含8000+道题,覆盖17个领域,最先进模型准确率仅86%。
- 适合研究跨语言视觉推理、公平考试准备及电子政务系统评估的学者和工程师。
我们提出EuraGovExam,一个源自五个欧亚代表性地区(韩国、日本、台湾、印度、欧盟)真实公务员考试的多语言多模态基准。该数据集包含超过8,000张高分辨率扫描的多选题,覆盖17个多样化的学术与行政领域。不同于现有基准,EuraGovExam将问题陈述、答案选项与视觉元素全部嵌入单张图像中,仅提供最小化统一的答案格式指令。这一设计要求模型在视觉输入上执行布局感知与跨语言推理。所有题目均来自真实考试文档,保留了表格、多语言排版和表单式布局等丰富视觉结构。评估结果显示,即使是最先进的视觉语言模型(VLMs)也仅达到86%准确率,凸显该基准的难度及其诊断当前模型局限性的能力。通过强调文化真实性、视觉复杂性与语言多样性,EuraGovExam为高风险、多语言、图像驱动场景下的VLM评估设立了新标准。同时支持电子政务、公共部门文档分析与公平备考等实际应用。
原文摘要 · Abstract (English)
We present EuraGovExam, a multilingual and multimodal benchmark sourced from real-world civil service examinations across five representative Eurasian regions: South Korea, Japan, Taiwan, India, and the European Union. Designed to reflect the authentic complexity of public-sector assessments, the dataset contains over 8,000 high-resolution scanned multiple-choice questions covering 17 diverse academic and administrative domains. Unlike existing benchmarks, EuraGovExam embeds all question content--including problem statements, answer choices, and visual elements--within a single image, providing only a minimal standardized instruction for answer formatting. This design demands that models perform layout-aware, cross-lingual reasoning directly from visual input. All items are drawn from real exam documents, preserving rich visual structures such as tables, multilingual typography, and form-like layouts. Evaluation results show that even state-of-the-art vision-language models (VLMs) achieve only 86% accuracy, underscoring the benchmark's difficulty and its power to diagnose the limitations of current models. By emphasizing cultural realism, visual complexity, and linguistic diversity, EuraGovExam establishes a new standard for evaluating VLMs in high-stakes, multilingual, image-grounded settings. It also supports practical applications in e-governance, public-sector document analysis, and equitable exam preparation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。