首个库尔德语索拉尼版命名实体识别数据集,挑战神经模型在低资源语言中的优势。
Named Entity Recognition for the Kurdish Sorani Language: Dataset Creation and Comparative Analysis
- 构建首个索拉尼语命名实体识别数据集,含64,563个标注词元
- 传统CRF模型F1达0.825,显著优于BiLSTM的0.706
- 适合低资源语言研究者、NLP公平性关注者参考
本研究致力于提升自然语言处理技术的包容性与全球适用性,提出首个针对库尔德语索拉尼语的命名实体识别数据集,包含64,563个标注词元。同时提供一款支持该语言及其他多种语言的标注工具,并开展全面的对比分析,涵盖经典机器学习模型与神经网络系统。结果表明,在低资源场景下,传统方法具有显著优势:条件随机场(CRF)模型获得0.825的F1分数,明显高于基于BiLSTM的模型(0.706)。这一发现挑战了神经方法在自然语言处理中普遍占优的既有认知,说明在资源有限的情况下,更简单、计算效率更高的经典框架仍具竞争力。
原文摘要 · Abstract (English)
This work contributes towards balancing the inclusivity and global applicability of natural language processing techniques by proposing the first 'name entity recognition' dataset for Kurdish Sorani, a low-resource and under-represented language, that consists of 64,563 annotated tokens. It also provides a tool for facilitating this task in this and many other languages and performs a thorough comparative analysis, including classic machine learning models and neural systems. The results obtained challenge established assumptions about the advantage of neural approaches within the context of NLP. Conventional methods, in particular CRF, obtain F1-scores of 0.825, outperforming the results of BiLSTM-based models (0.706) significantly. These findings indicate that simpler and more computationally efficient classical frameworks can outperform neural architectures in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。