arXiv:2410.20238cs.CLcs.AI2024-10综述被引 33

系统梳理阿拉伯语大模型研究现状与挑战

A Survey of Large Language Models for Arabic Language and its Dialects

  • 分类总结阿拉伯语各类大模型架构与训练数据
  • 评估模型开源程度,强调透明性对研究重要性
  • 指出方言数据稀缺,适合语言学与AI交叉研究者

本综述全面回顾了面向阿拉伯语及其方言的大型语言模型(LLMs)。涵盖编码器-仅用、解码器-仅用及编码器-解码器等关键架构,以及用于预训练的数据集,包括古典阿拉伯语、现代标准阿拉伯语和方言阿拉伯语。研究还探讨了单语、双语和多语种大模型,分析其在情感分析、命名实体识别和问答等下游任务中的表现。同时,基于源代码、训练数据、模型权重和文档的可获取性,评估阿拉伯语大模型的开放程度。综述强调需构建更多元化的方言数据集,并指出开放性对研究可复现性和透明性的关键作用。最后,识别出未来研究的关键挑战与机遇,呼吁开发更具包容性和代表性的模型。

原文摘要 · Abstract (English)

This survey offers a comprehensive overview of Large Language Models (LLMs) designed for Arabic language and its dialects. It covers key architectures, including encoder-only, decoder-only, and encoder-decoder models, along with the datasets used for pre-training, spanning Classical Arabic, Modern Standard Arabic, and Dialectal Arabic. The study also explores monolingual, bilingual, and multilingual LLMs, analyzing their architectures and performance across downstream tasks, such as sentiment analysis, named entity recognition, and question answering. Furthermore, it assesses the openness of Arabic LLMs based on factors, such as source code availability, training data, model weights, and documentation. The survey highlights the need for more diverse dialectal datasets and attributes the importance of openness for research reproducibility and transparency. It concludes by identifying key challenges and opportunities for future research and stressing the need for more inclusive and representative models.

大模型阿拉伯语语言技术开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。