构建首个覆盖东南亚多语言文化的LLM评估体系
SEA-HELM: Southeast Asian Holistic Evaluation of Language Models
- 设计五大维度评估框架,涵盖语言、文化与安全
- 支持菲律宾语、印尼语等5种东南亚语言评测
- 适合关注跨文化AI公平性的研究者与开发者
随着大语言模型(LLMs)新能力的快速涌现,对多语言、多文化整合评估基准的需求愈发迫切。尽管现有基准可评估英语及部分中低资源语言(包括东南亚地区语言)的特定能力,但尚未有全面且具文化代表性的东南亚语言评估体系。本文提出SEA-HELM,一个聚焦东南亚语言的综合性语言与文化评估套件,包含五大核心支柱:(1) NLP经典任务,(2) LLM专用任务,(3) 东南亚语言学特性,(4) 东南亚文化理解,(5) 安全性。当前支持菲律宾语、印尼语、泰米尔语、泰语和越南语。我们还推出了SEA-HELM排行榜,以系统化、用户友好的方式展示模型在多语言与多文化场景下的表现。评估代码已公开可用。
原文摘要 · Abstract (English)
With the rapid emergence of novel capabilities in Large Language Models (LLMs), the need for rigorous multilingual and multicultural benchmarks that are integrated has become more pronounced. Though existing LLM benchmarks are capable of evaluating specific capabilities of LLMs in English as well as in various mid- to low-resource languages, including those in the Southeast Asian (SEA) region, a comprehensive and culturally representative evaluation suite for the SEA languages has not been developed thus far. Here, we present SEA-HELM, a holistic linguistic and cultural LLM evaluation suite that emphasises SEA languages, comprising five core pillars: (1) NLP Classics, (2) LLM-specifics, (3) SEA Linguistics, (4) SEA Culture, (5) Safety. SEA-HELM currently supports Filipino, Indonesian, Tamil, Thai, and Vietnamese. We also introduce the SEA-HELM leaderboard, which allows users to understand models' multilingual and multicultural performance in a systematic and user-friendly manner. We make the SEA-HELM evaluation code publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。