arXiv:2507.08924cs.CLcs.AI2025-07EMNLP被引 5

构建韩语专业级大模型评测基准,覆盖真实职业场景。

From KMMLU-Redux to KMMLU-Pro: A Professional Korean Benchmark Suite for LLM Evaluation

  • 基于韩国职业资格考试重构评测题库,提升可靠性。
  • 涵盖韩国国家级专业技术考试内容,反映工业知识真实水平。
  • 公开数据集,适合评估大模型在韩语专业场景表现。

大语言模型的发展需要涵盖学术与产业领域的可靠评测基准,以有效评估其在真实场景中的适用性。本文提出两个韩语专家级评测基准:KMMLU-Redux 是对现有 KMMLU 的重构,题源来自韩国国家技术资格考试,剔除了关键错误以增强可靠性;KMMLU-Pro 基于韩国国家专业执照考试,反映韩国真实职业知识体系。实验表明,这些基准能全面代表韩国产业知识。数据集已公开发布。

原文摘要 · Abstract (English)

The development of Large Language Models (LLMs) requires robust benchmarks that encompass not only academic domains but also industrial fields to effectively evaluate their applicability in real-world scenarios. In this paper, we introduce two Korean expert-level benchmarks. KMMLU-Redux, reconstructed from the existing KMMLU, consists of questions from the Korean National Technical Qualification exams, with critical errors removed to enhance reliability. KMMLU-Pro is based on Korean National Professional Licensure exams to reflect professional knowledge in Korea. Our experiments demonstrate that these benchmarks comprehensively represent industrial knowledge in Korea. We release our dataset publicly available.

大模型评测韩语职业知识基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。