首个藏语理解评估基准,填补大模型在藏语上的评估空白。
TLUE: A Tibetan Language Understanding Evaluation Benchmark
- 构建涵盖5大领域67子任务的多任务藏语理解基准
- 多数大模型表现低于随机水平,揭示藏语处理重大挑战
- 适合关注低资源语言、模型公平性研究者使用
近年来大语言模型发展迅速,但低资源语言如藏语在评估中仍严重缺失。尽管藏语使用者超过七百万,却在大模型研发与评估中长期被忽视。为此,我们提出首个大规模藏语理解评估基准TLUE,包含覆盖5个领域、67个子任务的综合性多任务理解基准,以及涵盖7个子任务的安全性基准。我们对一系列先进大语言模型进行了评估,实验结果表明,大多数模型表现低于随机基线,凸显其在藏语处理上的巨大挑战。TLUE为未来藏语理解研究提供了关键基础,并强调了提升大模型开发包容性的必要性。
原文摘要 · Abstract (English)
Large language models have made tremendous progress in recent years, but low-resource languages, like Tibetan, remain significantly underrepresented in their evaluation. Despite Tibetan being spoken by over seven million people, it has largely been neglected in the development and assessment of large language models. To address this gap, we present a \textbf{T}ibetan \textbf{L}anguage \textbf{U}nderstanding \textbf{E}valuation Benchmark, \textbf{TLUE}, the first large-scale benchmark for measuring the proficiency of LLMs in the Tibetan language. \textbf{TLUE} comprises two major components: a comprehensive multi-task understanding benchmark spanning 5 domains and 67 subdomains, and a safety benchmark encompassing 7 subdomains. Then, we evaluate a diverse set of state-of-the-art large language models. Experimental results demonstrate that most large language models perform below the random baseline, highlighting the considerable challenges they face in Tibetan language processing. \textbf{TLUE} provides a crucial foundation for advancing future research in Tibetan language understanding and highlights the importance of promoting greater inclusivity in the development of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。