arXiv:2506.19753cs.CLcs.AI2025-06被引 2

对比多种模型,提升阿拉伯语方言分类准确率

Arabic Dialect Classification using RNNs, Transformers, and Large Language Models: A Comparative Analysis

  • 采用RNN、Transformer与大模型结合提示工程进行分类
  • MARBERTv2在18种方言上达65%准确率与64%F1分数
  • 适合做方言个性化对话系统与社交媒体分析

阿拉伯语是全球使用最广泛的语言之一,涵盖22个国家的众多方言。本文针对阿拉伯推文数据集QADI中的18种阿拉伯语方言分类问题,构建并测试了RNN模型、Transformer模型以及通过提示工程的大语言模型(LLMs)。其中,MARBERTv2表现最佳,达到65%的准确率和64%的F1分数。结合最先进的预处理技术和最新NLP模型,本研究识别出阿拉伯语方言识别中的关键语言挑战。结果支持个性化聊天机器人、社交媒体监控及阿拉伯语社区可访问性等应用。

原文摘要 · Abstract (English)

The Arabic language is among the most popular languages in the world with a huge variety of dialects spoken in 22 countries. In this study, we address the problem of classifying 18 Arabic dialects of the QADI dataset of Arabic tweets. RNN models, Transformer models, and large language models (LLMs) via prompt engineering are created and tested. Among these, MARBERTv2 performed best with 65% accuracy and 64% F1-score. Through the use of state-of-the-art preprocessing techniques and the latest NLP models, this paper identifies the most significant linguistic issues in Arabic dialect identification. The results corroborate applications like personalized chatbots that respond in users' dialects, social media monitoring, and greater accessibility for Arabic communities.

方言识别大模型自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。