arXiv:2503.22968cs.CEcs.AI2025-03中稿 · LREC 2026被引 1

构建统一评估框架,解决韩语大模型评测标准不一问题

Redefining Evaluation Standards: A Unified Framework for Evaluating the Korean Capabilities of Language Models

  • 基于注册表设计的开源工具HRET整合主流韩语评测集与多推理后端
  • 引入语法一致性和关键词遗漏检测等专有分析,揭示模型输出缺陷
  • 支持快速集成新数据集与方法,适合韩语大模型研发与评测人员

近期韩语大语言模型的发展催生了众多评测基准与方法,但协议不一致导致机构间性能差距高达10个百分点。解决可复现性问题并非要求统一评测模板,而是需要支持多样化实验的稳健框架。为此,我们提出HRET(Haerae评估工具包),一个基于注册表的开源框架,统一韩语大模型评估。HRET集成主要韩语评测集、多种推理后端及多方法评估,并通过语言一致性约束确保真实韩语输出。其模块化注册表设计支持快速引入新数据集、方法和后端,适应研究演进需求。除标准准确率外,还包含针对韩语的诊断分析:考虑词形的类型-词汇比(TTR)用于评估词汇多样性,系统性关键词遗漏检测用于识别缺失概念,帮助研究人员定位模型在形态与语义上的不足,指导韩语大模型的针对性优化。

原文摘要 · Abstract (English)

Recent advancements in Korean large language models (LLMs) have driven numerous benchmarks and evaluation methods, yet inconsistent protocols cause up to 10 p.p performance gaps across institutions. Overcoming these reproducibility gaps does not mean enforcing a one-size-fits-all evaluation. Rather, effective benchmarking requires diverse experimental approaches and a framework robust enough to support them. To this end, we introduce HRET (Haerae Evaluation Toolkit), an open-source, registry-based framework that unifies Korean LLM assessment. HRET integrates major Korean benchmarks, multiple inference backends, and multi-method evaluation, with language consistency enforcement to ensure genuine Korean outputs. Its modular registry design also enables rapid incorporation of new datasets, methods, and backends, ensuring the toolkit adapts to evolving research needs. Beyond standard accuracy metrics, HRET incorporates Korean-focused output analyses-morphology-aware Type-Token Ratio (TTR) for evaluating lexical diversity and systematic keyword-omission detection for identifying missing concepts-to provide diagnostic insights into language-specific behaviors. These targeted analyses help researchers pinpoint morphological and semantic shortcomings in model outputs, guiding focused improvements in Korean LLM development.

韩语模型评测框架语言评估大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。