KoBALT是评估韩语大模型语言理解力的全新基准,覆盖五大语言领域700道题。
KoBALT: Korean Benchmark For Advanced Linguistic Tasks
- 设计700道多选题,覆盖语法、语义、语用等5大语言领域,避免与常用语料重复。
- 顶尖模型在韩语理解上仅达61%准确率,语音和形态学表现最弱,分别仅31%和36%。
- 95名标注员验证显示,模型得分与人类判断高度一致,适合评估真实语言能力。
我们提出KoBALT(韩语高级语言任务基准),一个涵盖24种语言现象、5个语言领域(句法、语义、语用、音系/音位、形态)的综合性基准,包含700道多项选择题。该基准旨在推动韩语大语言模型(LLMs)的评估,解决传统基准缺乏语言深度与类型学基础的问题。所有题目均由专家设计,具有语言学动机,且与标准韩语语料的n-gram重叠极低,显著降低数据污染风险,实现对真正语言理解能力的稳健评估。对20个主流大模型的评测显示,最高性能模型总体准确率为61%,但在不同语言领域间差异显著:语义表现最佳(66%),而音系(31%)与形态(36%)表现最弱。通过95名标注员的人类偏好评估,我们验证了KoBALT得分与人类判断高度相关,证明其作为韩语理解判别性度量的有效性。KoBALT填补了类型多样语言评估的关键空白,为评估韩语模型的真实语言能力提供了可靠框架。
原文摘要 · Abstract (English)
We introduce KoBALT (Korean Benchmark for Advanced Linguistic Tasks), a comprehensive linguistically-motivated benchmark comprising 700 multiple-choice questions spanning 24 phenomena across five linguistic domains: syntax, semantics, pragmatics, phonetics/phonology, and morphology. KoBALT is designed to advance the evaluation of large language models (LLMs) in Korean, a morphologically rich language, by addressing the limitations of conventional benchmarks that often lack linguistic depth and typological grounding. It introduces a suite of expert-curated, linguistically motivated questions with minimal n-gram overlap with standard Korean corpora, substantially mitigating the risk of data contamination and allowing a more robust assessment of true language understanding. Our evaluation of 20 contemporary LLMs reveals significant performance disparities, with the highest-performing model achieving 61\% general accuracy but showing substantial variation across linguistic domains - from stronger performance in semantics (66\%) to considerable weaknesses in phonology (31\%) and morphology (36\%). Through human preference evaluation with 95 annotators, we demonstrate a strong correlation between KoBALT scores and human judgments, validating our benchmark's effectiveness as a discriminative measure of Korean language understanding. KoBALT addresses critical gaps in linguistic evaluation for typologically diverse languages and provides a robust framework for assessing genuine linguistic competence in Korean language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。