arXiv:2510.15558cs.CLcs.AI2025-10

首个评估大模型韩语指令遵循能力的基准,填补语言文化空白。

KITE: A Benchmark for Evaluating Korean Instruction-Following Abilities in Large Language Models

  • 构建韩语专属开放指令任务评测集,覆盖语法与文化特性。
  • 结合自动评分与人工评估,发现不同模型在韩语任务中表现差异显著。
  • 适合关注多语言AI公平性与跨文化模型研发的研究者。

大型语言模型(LLM)的指令遵循能力对对话系统、复杂推理等应用至关重要。然而,当前评估主要聚焦英文模型,忽视了其他语言的语言与文化特性。韩语具有独特的语法结构、丰富的形态特征、敬语体系及双重数词系统,却缺乏针对开放式指令遵循能力的专用评测基准。为此,我们提出韩国指令遵循任务评估基准(KITE),全面评估通用与韩语特有指令任务。与以往侧重事实知识或选择题的韩语评测不同,KITE直接面向多样化的开放指令任务。评估流程融合自动化指标与人工评价,揭示模型间性能差异,并深入分析其优劣势。通过公开发布KITE数据集与代码,我们旨在推动更具文化与语言包容性的大模型研究,激励其他低资源语言开展类似工作。

原文摘要 · Abstract (English)

The instruction-following capabilities of large language models (LLMs) are pivotal for numerous applications, from conversational agents to complex reasoning systems. However, current evaluations predominantly focus on English models, neglecting the linguistic and cultural nuances of other languages. Specifically, Korean, with its distinct syntax, rich morphological features, honorific system, and dual numbering systems, lacks a dedicated benchmark for assessing open-ended instruction-following capabilities. To address this gap, we introduce the Korean Instruction-following Task Evaluation (KITE), a comprehensive benchmark designed to evaluate both general and Korean-specific instructions. Unlike existing Korean benchmarks that focus mainly on factual knowledge or multiple-choice testing, KITE directly targets diverse, open-ended instruction-following tasks. Our evaluation pipeline combines automated metrics with human assessments, revealing performance disparities across models and providing deeper insights into their strengths and weaknesses. By publicly releasing the KITE dataset and code, we aim to foster further research on culturally and linguistically inclusive LLM development and inspire similar endeavors for other underrepresented languages.

指令遵循韩语评测基准多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。