arXiv:2508.12459cs.CL2025-08EMNLP被引 4

构建首个覆盖20种印尼语的多任务多语言基准,推动低资源语言NLP发展

LoraxBench: A Multitask, Multilingual Benchmark Suite for 20 Indonesian Languages

  • 设计涵盖6类任务的多语言评估体系,覆盖20种印尼语及3种语言的正式度变体
  • 实测发现印尼语模型性能显著优于其他低资源语言,且区域模型无明显优势
  • 揭示正式度差异(如爪哇语'Krama')对模型表现有显著影响,尤其在非社交媒体语境

作为全球人口最多的国家之一,印尼拥有700种方言,但自然语言处理进展滞后。本文提出LoraxBench,一个聚焦印尼低资源语言的多任务、多语言基准,涵盖阅读理解、开放域问答、语言推理、因果推理、翻译和文化问答共6项任务。数据集覆盖20种语言,并为其中3种语言增加两种正式度变体。我们评估了多种多语言与区域专用大模型,发现该基准极具挑战性:印尼语模型表现远超其他语言,尤其是低资源语言;区域模型与通用多语言模型之间未见明显优劣。此外,正式度变化显著影响模型性能,尤其在非社交媒体语境中常见但罕见于训练数据的高礼貌形式(如爪哇语Krama)。

原文摘要 · Abstract (English)

As one of the world's most populous countries, with 700 languages spoken, Indonesia is behind in terms of NLP progress. We introduce LoraxBench, a benchmark that focuses on low-resource languages of Indonesia and covers 6 diverse tasks: reading comprehension, open-domain QA, language inference, causal reasoning, translation, and cultural QA. Our dataset covers 20 languages, with the addition of two formality registers for three languages. We evaluate a diverse set of multilingual and region-focused LLMs and found that this benchmark is challenging. We note a visible discrepancy between performance in Indonesian and other languages, especially the low-resource ones. There is no clear lead when using a region-specific model as opposed to the general multilingual model. Lastly, we show that a change in register affects model performance, especially with registers not commonly found in social media, such as high-level politeness `Krama' Javanese.

多语言低资源基准测试印尼语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。