构建韩语语法评估基准,揭示大模型在真实语用知识上的短板
Polishing Every Facet of the GEM: Testing Linguistic Competence of LLMs and Humans in Korean
- 设计1.5千道多选题的韩语语法评测集,覆盖5大类16子类
- 27个大模型零样本测试显示:定义性任务表现好,语用知识整合差
- 发现语音规则等实感知识是提升语言理解的关键,适合研究者与开发者
我们提出韩国语法评估基准(KoGEM),用于评估大语言模型与人类在韩语中的语言能力。KoGEM包含1500个多项选择问答对,涵盖五大类别与16个子类别。对27种不同规模和类型的大型语言模型进行零样本评估发现,尽管它们在依赖定义性知识的简单任务中表现优异,但在需要整合现实世界经验知识的任务(如语音规则与发音)上表现不佳。深入分析表明,融入此类经验知识可显著提升大模型的语言能力。KoGEM不仅揭示了当前大模型在语言理解上的局限性,也挖掘出其潜在的未被充分开发的语言认知维度,为全面提升语言理解能力提供方向。代码与数据集已公开于:https://github.com/SungHo3268/KoGEM。
原文摘要 · Abstract (English)
We introduce the $\underline{Ko}rean \underline{G}rammar \underline{E}valuation Bench\underline{M}ark (KoGEM)$, designed to assess the linguistic competence of LLMs and humans in Korean. KoGEM consists of 1.5k multiple-choice QA pairs covering five main categories and 16 subcategories. The zero-shot evaluation of 27 LLMs of various sizes and types reveals that while LLMs perform remarkably well on straightforward tasks requiring primarily definitional knowledge, they struggle with tasks that demand the integration of real-world experiential knowledge, such as phonological rules and pronunciation. Furthermore, our in-depth analysis suggests that incorporating such experiential knowledge could enhance the linguistic competence of LLMs. With KoGEM, we not only highlight the limitations of current LLMs in linguistic competence but also uncover hidden facets of LLMs in linguistic competence, paving the way for enhancing comprehensive language understanding. Our code and dataset are available at: https://github.com/SungHo3268/KoGEM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。