针对低资源语言理解难题,构建了文化特异性谚语评估基准。
ProverbEval: Exploring LLM Evaluation Challenges for Low-resource Language Understanding
- 设计文化特异性谚语任务,评估大模型在低资源语言中的理解能力。
- 答案选项顺序差异导致性能波动高达50%,提示评估设计需谨慎。
- 使用母语谚语描述可显著提升生成任务表现,适合跨语言研究者参考。
随着评估数据集的快速发展,为大模型在多领域、多主题下的理解能力设定合适的评测基准变得愈发困难。本文聚焦低资源语言理解中的评估挑战,提出 ProverbEval,一个面向低资源语言文化特异性场景的大模型评估基准。我们对多种大模型进行了评测,并探究了评测过程中影响结果变异的关键因素。结果显示,多项选择题中答案选项的呈现顺序会导致高达50%的性能波动。使用母语谚语描述能显著提升谚语生成等任务的表现。此外,在生成任务中,单语评估始终优于跨语言评估。我们强调在构建大模型评估基准时,必须关注选项顺序、提示语言选择、任务变异性及生成任务设计。评测数据见 https://huggingface.co/datasets/israel/ProverbEval,代码开源于 https://github.com/EthioNLP/EthioProverbEval。
原文摘要 · Abstract (English)
With the rapid development of evaluation datasets to assess LLMs understanding across a wide range of subjects and domains, identifying a suitable language understanding benchmark has become increasingly challenging. In this work, we explore LLM evaluation challenges for low-resource language understanding and introduce \proverbeval, LLM evaluation benchmark for low-resource languages, focusing on low-resource language understanding in culture-specific scenarios. We benchmark various LLMs and explore factors that create variability in the benchmarking process. We observed performance variances of up to 50\%, depending on the order in which answer choices were presented in multiple-choice tasks. Native language proverb descriptions significantly improve tasks such as proverb generation, contributing to improved outcomes. Additionally, monolingual evaluations consistently outperformed their cross-lingual counterparts in generation tasks. We argue that special attention must be given to the order of choices, the choice of prompt language, task variability, and generation tasks when creating LLM evaluation benchmarks. Evaluation data available at https://huggingface.co/datasets/israel/ProverbEval, evaluation code https://github.com/EthioNLP/EthioProverbEval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。