arXiv:2502.15422cs.CLcs.AI2025-02NAACL被引 7

构建韩语教育测评基准,评估多模态生成AI在不同学段的表现

Evaluating Multimodal Generative AI with Korean Educational Standards

  • 基于韩国四大教育考试构建评测基准,覆盖小初高及大学水平
  • 首次系统评估多模态AI在韩语环境下的答题能力与错误率
  • 开源代码与数据构建工具,助力非英语语言AI研究

本文提出韩国国家教育测试基准(KoNET),用于评估多模态生成式AI系统在韩语教育考试中的表现。KoNET包含四项考试:小学通用教育发展测试(KoEGED)、中学通用教育发展测试(KoMGED)、高中通用教育发展测试(KoHGED)和大学修业能力测试(KoCSAT)。这些考试以高标准和多样化题型著称,可全面分析AI在不同教育阶段的表现。通过聚焦韩语,KoNET为少研究语言的模型性能提供了新视角。我们评估了多种模型——开源、开放访问及闭源API——涵盖难度、学科多样性及人类错误率。代码与数据集构建工具将完全开源,地址为 https://github.com/naver-ai/KoNET。

原文摘要 · Abstract (English)

This paper presents the Korean National Educational Test Benchmark (KoNET), a new benchmark designed to evaluate Multimodal Generative AI Systems using Korean national educational tests. KoNET comprises four exams: the Korean Elementary General Educational Development Test (KoEGED), Middle (KoMGED), High (KoHGED), and College Scholastic Ability Test (KoCSAT). These exams are renowned for their rigorous standards and diverse questions, facilitating a comprehensive analysis of AI performance across different educational levels. By focusing on Korean, KoNET provides insights into model performance in less-explored languages. We assess a range of models - open-source, open-access, and closed APIs - by examining difficulties, subject diversity, and human error rates. The code and dataset builder will be made fully open-sourced at https://github.com/naver-ai/KoNET.

多模态生成教育评测韩语AI基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。