arXiv:2504.21635cs.CLcs.AI2025-04被引 3

小模型Sadeed提升阿拉伯语拼音化准确率

Sadeed: Advancing Arabic Diacritization Through Small Language Model

  • 用微调的轻量语言模型处理阿拉伯语拼音
  • 在少量资源下性能媲美大模型,优于传统方法
  • 新基准SadeedDiac-25推动公平评估

阿拉伯语拼音化因词形复杂仍是自然语言处理难题。本文提出Sadeed,基于经过微调的解码器仅语言模型,该模型源自Kuwain 1.5B(Hennara et al. [2025]),原为多样化阿拉伯语语料训练的紧凑模型。Sadeed在精心清洗、规范化后的高质量拼音数据集上进行微调。尽管计算资源有限,其性能仍可与专有大模型竞争,并超越同类领域训练的传统模型。此外,我们指出现有阿拉伯语拼音化评测标准的关键缺陷。为此,我们构建SadeedDiac-25,一个新基准,旨在实现跨多种文本体裁和复杂度层级的更公平、全面评估。Sadeed与SadeedDiac-25共同为机器翻译、语音合成及语言学习工具等阿拉伯语NLP应用提供坚实基础。

原文摘要 · Abstract (English)

Arabic text diacritization remains a persistent challenge in natural language processing due to the language's morphological richness. In this paper, we introduce Sadeed, a novel approach based on a fine-tuned decoder-only language model adapted from Kuwain 1.5B Hennara et al. [2025], a compact model originally trained on diverse Arabic corpora. Sadeed is fine-tuned on carefully curated, high-quality diacritized datasets, constructed through a rigorous data-cleaning and normalization pipeline. Despite utilizing modest computational resources, Sadeed achieves competitive results compared to proprietary large language models and outperforms traditional models trained on similar domains. Additionally, we highlight key limitations in current benchmarking practices for Arabic diacritization. To address these issues, we introduce SadeedDiac-25, a new benchmark designed to enable fairer and more comprehensive evaluation across diverse text genres and complexity levels. Together, Sadeed and SadeedDiac-25 provide a robust foundation for advancing Arabic NLP applications, including machine translation, text-to-speech, and language learning tools.

阿拉伯语拼音化小模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。