arXiv:2501.14528cs.CL2025-01被引 6

首个针对索拉尼库尔德语的习语检测研究,用深度学习识别隐喻表达。

Idiom Detection in Sorani Kurdish Texts

  • 将习语检测建模为文本分类,采用Transformer、RCNN和BiLSTM三种模型
  • KuBERT模型达到99%准确率,显著优于RCNN(96.5%)和BiLSTM(80%)
  • 开源10,580句带101个习语的数据集,助力低资源语言NLP发展

利用自然语言处理进行习语检测是计算机识别文本中超越字面意义的修辞表达的过程。尽管多种语言在该领域已取得进展,但库尔德语在这一方向仍存在明显研究空白,而习语在机器翻译和情感分析等任务中至关重要。本研究针对索拉尼库尔德语开展习语检测,将其作为文本分类任务,采用深度学习方法。为此,构建了一个包含10,580个句子、涵盖101个索拉尼库尔德语习语的语料库。基于此数据集,开发并评估了三种模型:基于KuBERT的Transformer序列分类模型、循环卷积神经网络(RCNN)以及带注意力机制的BiLSTM模型。实验结果表明,微调后的Transformer模型持续表现最优,准确率达99%;RCNN为96.5%;BiLSTM为80%。这些结果凸显了Transformer架构在低资源语言如库尔德语中的有效性。本研究提供了数据集、三个优化模型及检测见解,为推进库尔德语NLP奠定基础。

原文摘要 · Abstract (English)

Idiom detection using Natural Language Processing (NLP) is the computerized process of recognizing figurative expressions within a text that convey meanings beyond the literal interpretation of the words. While idiom detection has seen significant progress across various languages, the Kurdish language faces a considerable research gap in this area despite the importance of idioms in tasks like machine translation and sentiment analysis. This study addresses idiom detection in Sorani Kurdish by approaching it as a text classification task using deep learning techniques. To tackle this, we developed a dataset containing 10,580 sentences embedding 101 Sorani Kurdish idioms across diverse contexts. Using this dataset, we developed and evaluated three deep learning models: KuBERT-based transformer sequence classification, a Recurrent Convolutional Neural Network (RCNN), and a BiLSTM model with an attention mechanism. The evaluations revealed that the transformer model, the fine-tuned BERT, consistently outperformed the others, achieving nearly 99% accuracy while the RCNN achieved 96.5% and the BiLSTM 80%. These results highlight the effectiveness of Transformer-based architectures in low-resource languages like Kurdish. This research provides a dataset, three optimized models, and insights into idiom detection, laying a foundation for advancing Kurdish NLP.

习语检测库尔德语深度学习低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。