arXiv:2508.19966cs.CLcs.AI2025-08

用自建数据集微调大模型,实现97.8%准确率的阿拉伯语主观性分析

Dhati+: Fine-tuned Large Language Models for Arabic Subjectivity Evaluation

  • 构建新数据集AraDhati+,融合多个现有阿拉伯语语料
  • 在AraDhati+上微调XLM-RoBERTa等模型,达97.79%准确率
  • 适合从事阿拉伯语文本分析、低资源语言处理的研究者

尽管重要,阿拉伯语因语言丰富且形态复杂,仍面临资源匮乏问题。缺乏大规模标注数据制约了其主观性分析工具的发展。近年来深度学习与Transformer模型在英语和法语文本分类中表现优异。本文提出一种新的阿拉伯语主观性评估方法:为弥补专用标注数据不足,我们利用现有阿拉伯语数据集(ASTD、LABR、HARD和SANAD)构建了综合性数据集AraDhati+;随后在AraDhati+上对XLM-RoBERTa、AraBERT和ArabianGPT等先进阿拉伯语模型进行微调,以实现有效的主观性分类;此外,还尝试了集成决策方法以融合各模型优势。实验表明,该方法在阿拉伯语主观性分类任务上达到97.79%的惊人准确率,验证了其在应对阿拉伯语资源有限挑战方面的有效性。

原文摘要 · Abstract (English)

Despite its significance, Arabic, a linguistically rich and morphologically complex language, faces the challenge of being under-resourced. The scarcity of large annotated datasets hampers the development of accurate tools for subjectivity analysis in Arabic. Recent advances in deep learning and Transformers have proven highly effective for text classification in English and French. This paper proposes a new approach for subjectivity assessment in Arabic textual data. To address the dearth of specialized annotated datasets, we developed a comprehensive dataset, AraDhati+, by leveraging existing Arabic datasets and collections (ASTD, LABR, HARD, and SANAD). Subsequently, we fine-tuned state-of-the-art Arabic language models (XLM-RoBERTa, AraBERT, and ArabianGPT) on AraDhati+ for effective subjectivity classification. Furthermore, we experimented with an ensemble decision approach to harness the strengths of individual models. Our approach achieves a remarkable accuracy of 97.79\,\% for Arabic subjectivity classification. Results demonstrate the effectiveness of the proposed approach in addressing the challenges posed by limited resources in Arabic language processing.

阿拉伯语主观性分析大模型微调文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。