首个针对濒危南岛语的LLM评测基准,揭示大模型在少数民族语言上的显著性能差距。
FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models
- 构建首个面向台湾南岛语的多任务评测基准,覆盖3种濒危语言
- 零样本与少量样本下大模型表现远低于高资源语言,微调提升有限
- 适合关注语言多样性、低资源NLP及濒危语言保护的研究者
尽管大型语言模型(LLMs)在高资源语言上表现出色,但在低资源和少数语言上的能力仍严重不足。台湾原住民语言——南岛语系的福尔摩沙语,语言丰富但面临消亡风险,主要因普通话的强势地位。本文提出FORMOSANBENCH,首个评估大模型在低资源南岛语上表现的基准,涵盖阿美语、泰雅语和排湾语三种濒危语言,覆盖机器翻译、自动语音识别(ASR)和文本摘要三个核心自然语言处理任务。我们在零样本、10样本和微调设置下评估模型性能。结果显示,高资源语言与福尔摩沙语之间存在显著性能差距;现有大模型在所有任务中均表现不佳,10样本学习和微调仅带来有限改进。这些发现凸显了开发更具包容性的自然语言技术以支持濒危语言的迫切需求。我们公开数据集与代码,以推动该方向的后续研究。
原文摘要 · Abstract (English)
While large language models (LLMs) have demonstrated impressive performance across a wide range of natural language processing (NLP) tasks in high-resource languages, their capabilities in low-resource and minority languages remain significantly underexplored. Formosan languages -- a subgroup of Austronesian languages spoken in Taiwan -- are both linguistically rich and endangered, largely due to the sociolinguistic dominance of Mandarin. In this work, we introduce FORMOSANBENCH, the first benchmark for evaluating LLMs on low-resource Austronesian languages. It covers three endangered Formosan languages: Atayal, Amis, and Paiwan, across three core NLP tasks: machine translation, automatic speech recognition (ASR), and text summarization. We assess model performance in zero-shot, 10-shot, and fine-tuned settings using FORMOSANBENCH. Our results reveal a substantial performance gap between high-resource and Formosan languages. Existing LLMs consistently underperform across all tasks, with 10-shot learning and fine-tuning offering only limited improvements. These findings underscore the urgent need for more inclusive NLP technologies that can effectively support endangered and underrepresented languages. We release our datasets and code to facilitate future research in this direction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。