评测大模型对突尼斯阿拉伯语的理解能力,揭示其语言鸿沟。
How Well Do LLMs Understand Tunisian Arabic?
- 构建含平行文本与情感标签的突尼斯阿拉伯语数据集
- 多模型在转写、翻译、情感分析任务上表现差异显著
- 推动低资源语言融入下一代AI,保障技术包容性
大型语言模型(LLMs)是当今人工智能代理的核心引擎。模型对人类语言的理解越深入,与AI的交互就越自然友好,涵盖从电脑、智能手表到任何具备智能功能的工具。然而,工业级大模型对低资源语言如突尼斯阿拉伯语(Tunizi)的理解能力常被忽视。这种忽视可能导致数百万突尼斯人无法用母语充分使用AI,被迫转向法语或英语,不仅威胁突尼斯方言的传承,还可能影响识字率,并使年轻一代更倾向于使用外语。本研究引入一个包含并行突尼斯阿拉伯语、标准突尼斯阿拉伯语和英文翻译及情感标签的新数据集。我们在三个任务上评估多个主流大模型:转写、翻译和情感分析。结果揭示了模型间显著差异,凸显其在理解和处理突尼斯方言时的优势与局限。通过量化这些差距,本工作强调将低资源语言纳入下一代AI系统的重要性,确保技术更具可访问性、包容性和文化根基。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are the engines driving today's AI agents. The better these models understand human languages, the more natural and user-friendly the interaction with AI becomes, from everyday devices like computers and smartwatches to any tool that can act intelligently. Yet, the ability of industrial-scale LLMs to comprehend low-resource languages, such as Tunisian Arabic (Tunizi), is often overlooked. This neglect risks excluding millions of Tunisians from fully interacting with AI in their own language, pushing them toward French or English. Such a shift not only threatens the preservation of the Tunisian dialect but may also create challenges for literacy and influence younger generations to favor foreign languages. In this study, we introduce a novel dataset containing parallel Tunizi, standard Tunisian Arabic, and English translations, along with sentiment labels. We benchmark several popular LLMs on three tasks: transliteration, translation, and sentiment analysis. Our results reveal significant differences between models, highlighting both their strengths and limitations in understanding and processing Tunisian dialects. By quantifying these gaps, this work underscores the importance of including low-resource languages in the next generation of AI systems, ensuring technology remains accessible, inclusive, and culturally grounded.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。