构建韩语教育测评基准,评估多模态生成AI在不同学段的表现
Evaluating Multimodal Generative AI with Korean Educational Standards
- 基于韩国四大教育考试构建评测基准,覆盖小初高及大学水平
- 首次系统评估多模态AI在韩语环境下的答题能力与错误率
- 开源代码与数据构建工具,助力非英语语言AI研究
本文提出韩国国家教育测试基准(KoNET),用于评估多模态生成式AI系统在韩语教育考试中的表现。KoNET包含四项考试:小学通用教育发展测试(KoEGED)、中学通用教育发展测试(KoMGED)、高中通用教育发展测试(KoHGED)和大学修业能力测试(KoCSAT)。这些考试以高标准和多样化题型著称,可全面分析AI在不同教育阶段的表现。通过聚焦韩语,KoNET为少研究语言的模型性能提供了新视角。我们评估了多种模型——开源、开放访问及闭源API——涵盖难度、学科多样性及人类错误率。代码与数据集构建工具将完全开源,地址为 https://github.com/naver-ai/KoNET。
原文摘要 · Abstract (English)
This paper presents the Korean National Educational Test Benchmark (KoNET), a new benchmark designed to evaluate Multimodal Generative AI Systems using Korean national educational tests. KoNET comprises four exams: the Korean Elementary General Educational Development Test (KoEGED), Middle (KoMGED), High (KoHGED), and College Scholastic Ability Test (KoCSAT). These exams are renowned for their rigorous standards and diverse questions, facilitating a comprehensive analysis of AI performance across different educational levels. By focusing on Korean, KoNET provides insights into model performance in less-explored languages. We assess a range of models - open-source, open-access, and closed APIs - by examining difficulties, subject diversity, and human error rates. The code and dataset builder will be made fully open-sourced at https://github.com/naver-ai/KoNET.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。