arXiv:2412.07251cs.CL2024-12中稿 · the 38th Pacific A…被引 8

专为韩语文化设计的评估基准,测试大模型对韩国文化的理解能力。

KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context

  • 构建韩语文化专用评测集,涵盖新闻、谚语和诗歌。
  • 多模型对比显示,现有模型对深层韩文化理解仍不充分。
  • 适合研究跨文化语言模型或韩语NLP的学者使用。

大型语言模型在各类任务中表现显著提升,但其生成内容日益流畅连贯,评估难度随之增加。当前多语言评测数据常基于英文翻译,可能带有西方文化偏见,难以准确评估其他语言与文化。为此,我们提出KULTURE Bench,一个专为韩国文化设计的评估框架,包含文化新闻、习语和诗歌等数据集,用于在词、句、段层面评估模型的文化理解与推理能力。利用该基准,我们评估了使用不同语料训练的模型,并进行了全面分析。结果表明,模型对韩国文化深层内涵的理解仍有巨大提升空间。

原文摘要 · Abstract (English)

Large language models have exhibited significant enhancements in performance across various tasks. However, the complexity of their evaluation increases as these models generate more fluent and coherent content. Current multilingual benchmarks often use translated English versions, which may incorporate Western cultural biases that do not accurately assess other languages and cultures. To address this research gap, we introduce KULTURE Bench, an evaluation framework specifically designed for Korean culture that features datasets of cultural news, idioms, and poetry. It is designed to assess language models' cultural comprehension and reasoning capabilities at the word, sentence, and paragraph levels. Using the KULTURE Bench, we assessed the capabilities of models trained with different language corpora and analyzed the results comprehensively. The results show that there is still significant room for improvement in the models' understanding of texts related to the deeper aspects of Korean culture.

文化理解韩语NLP评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。