构建首个韩语否定理解基准,揭示大模型在否定句上的性能下降问题。
Thunder-KoNUBench: A Corpus-Aligned Benchmark for Korean Negation Understanding
- 基于真实语料构建韩语否定理解评测集,反映实际语言分布
- 47个大模型测试显示否定使性能显著下降,模型规模与指令微调有影响
- 该基准可有效提升模型对否定及上下文的理解能力,适合韩语NLP研究者
尽管否定现象已被证明会挑战大型语言模型(LLMs)的性能,但针对韩语否定理解的评测基准仍十分稀缺。我们基于语料库分析了韩语否定现象,发现大模型在否定句中的表现明显下降。为此,我们提出了Thunder-KoNUBench,一个反映韩语否定现象实际分布的句子级否定理解评测基准。我们在该基准上评估了47个大语言模型,分析了模型规模和指令微调的影响,并进行了错误分析以深入理解模型行为。此外,我们进一步验证了在Thunder-KoNUBench上进行微调,能有效提升模型在韩语中对否定及更广泛上下文的理解能力。
原文摘要 · Abstract (English)
Although negation is known to challenge large language models (LLMs), benchmarks for evaluating negation understanding-especially in Korean-are scarce. We conduct a corpus-based analysis of Korean negation and show that LLM performance degrades under negation. We then introduce Thunder-KoNUBench, a sentence-level negation understanding benchmark that reflects the empirical distribution of Korean negation phenomena. Evaluating 47 LLMs on Thunder-KoNUBench, we analyze the effects of model size and instruction tuning, and perform error analysis to better understand model behavior. We further show that fine-tuning on Thunder-KoNUBench improves negation understanding and broader contextual comprehension in Korean.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。