构建首个覆盖印尼本土文字的多模态多语言基准,揭示当前AI对本土文字处理能力严重不足。
NusaAksara: A Multimodal and Multilingual Benchmark for Preserving Indonesian Indigenous Scripts
- 涵盖8种印尼本土文字、7种语言,含非Unicode的Lampung文
- 测试多种大模型与专用系统,多数在本地文字任务上表现接近零
- 适合关注低资源语言、跨文化NLP及数字人文研究者
印度尼西亚拥有丰富的语言与文字体系,但当前自然语言处理进展主要基于拉丁化文本。本文提出NusaAksara,一个面向印尼语的公开多模态多语言基准,包含原始文字与图像数据,覆盖图像分割、OCR、音译、翻译、语言识别等多样任务。数据由专家严格构建,涵盖8种文字、7种语言,包括未被Unicode支持的Lampung文。我们在多个模型上进行评估,包括GPT-4o、Llama 3.2、Aya 23等大语言与视觉语言模型,以及PP-OCR、LangID等专用系统,结果显示多数NLP技术无法有效处理印尼本土文字,多项任务性能接近零。
原文摘要 · Abstract (English)
Indonesia is rich in languages and scripts. However, most NLP progress has been made using romanized text. In this paper, we present NusaAksara, a novel public benchmark for Indonesian languages that includes their original scripts. Our benchmark covers both text and image modalities and encompasses diverse tasks such as image segmentation, OCR, transliteration, translation, and language identification. Our data is constructed by human experts through rigorous steps. NusaAksara covers 8 scripts across 7 languages, including low-resource languages not commonly seen in NLP benchmarks. Although unsupported by Unicode, the Lampung script is included in this dataset. We benchmark our data across several models, from LLMs and VLMs such as GPT-4o, Llama 3.2, and Aya 23 to task-specific systems such as PP-OCR and LangID, and show that most NLP technologies cannot handle Indonesia's local scripts, with many achieving near-zero performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。