arXiv:2410.04795cs.CLcs.AI2024-10被引 6

为泰语大模型构建文化与核心能力双基准,填补评估空白。

Representing the Under-Represented: Cultural and Core Capability Benchmarks for Developing Thai Large Language Models

  • 提出泰语专属的H6与文化语言智能双基准
  • 验证多语言模型在泰语上的能力短板
  • 开源数据集与代码,助力本土大模型发展

大语言模型(LLM)的快速发展凸显了对其核心能力(如推理、知识、常识)进行有效评估的重要性,催生了如H6等广泛使用的基准套件。然而,这些基准主要针对英语设计,缺乏对泰语等欠代表性语言的支持。同时,泰语大模型的发展不仅需提升核心能力,还需增强文化理解。为此,我们提出两个关键基准:泰国-六(Thai-H6)与泰国文化与语言智能基准(ThaiCLI)。通过全面评估具备多语言能力的多种大模型,我们系统分析了所提基准对泰语大模型发展的贡献。此外,我们将公开数据集与评估代码,以推动泰语大模型的持续研究与开发。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) has highlighted the need for robust evaluation frameworks that assess their core capabilities, such as reasoning, knowledge, and commonsense, leading to the inception of certain widely-used benchmark suites such as the H6 benchmark. However, these benchmark suites are primarily built for the English language, and there exists a lack thereof for under-represented languages, in terms of LLM development, such as Thai. On the other hand, developing LLMs for Thai should also include enhancing the cultural understanding as well as core capabilities. To address these dual challenge in Thai LLM research, we propose two key benchmarks: Thai-H6 and Thai Cultural and Linguistic Intelligence Benchmark (ThaiCLI). Through a thorough evaluation of various LLMs with multi-lingual capabilities, we provide a comprehensive analysis of the proposed benchmarks and how they contribute to Thai LLM development. Furthermore, we will make both the datasets and evaluation code publicly available to encourage further research and development for Thai LLMs.

大模型评估泰语文化理解基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。