arXiv:2506.18399cs.CL2025-06EMNLP被引 1

将阿拉伯语词形还原任务转化为分类问题,提升准确性和可解释性。

Lemmatization as a Classification Task: Results from Arabic across Multiple Genres

  • 把词形还原转为标签分类,使用语义聚类和机器翻译构建新标签集。
  • 在多领域数据集上表现优于传统模型,准确率显著提升。
  • 适合需要高精度与可解释性的阿拉伯语NLP研究者使用。

词形还原对阿拉伯语等形态丰富的语言至关重要,但现有工具因标准不一、语域覆盖有限而面临挑战。本文提出两种新方法,将词形还原建模为对词形-词性-释义(LPG)标签集的分类任务,利用机器翻译和语义聚类构建标签体系。同时发布一个涵盖多种语域的新型阿拉伯语词形还原测试集,并与现有数据集统一标注标准。实验评估了字符级序列到序列模型,其在词形预测上表现良好,但仅支持词形输出,不生成完整LPG标签,且易生成不合理形式。相比之下,分类与聚类方法更具鲁棒性与可解释性,确立了阿拉伯语词形还原的新基准。

原文摘要 · Abstract (English)

Lemmatization is crucial for NLP tasks in morphologically rich languages with ambiguous orthography like Arabic, but existing tools face challenges due to inconsistent standards and limited genre coverage. This paper introduces two novel approaches that frame lemmatization as classification into a Lemma-POS-Gloss (LPG) tagset, leveraging machine translation and semantic clustering. We also present a new Arabic lemmatization test set covering diverse genres, standardized alongside existing datasets. We evaluate character level sequence-to-sequence models, which perform competitively and offer complementary value, but are limited to lemma prediction (not LPG) and prone to hallucinating implausible forms. Our results show that classification and clustering yield more robust, interpretable outputs, setting new benchmarks for Arabic lemmatization.

词形还原阿拉伯语分类任务NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。