用6854篇胡拉米语文章训练分类模型,线性SVM准确率达96%。
Shifting from endangerment to rebirth in the Artificial Intelligence Age: An Ensemble Machine Learning Approach for Hawrami Text Classification
- 整合四种机器学习模型,基于人工标注数据评估分类效果
- 线性SVM在15类胡拉米语文本中达到96%准确率
- 为濒危语言数字化提供可复用的文本分类范式
胡拉米语是库尔德语的一种方言,因数据稀缺和使用者减少被列为濒危语言。自然语言处理项目可通过机器翻译、语言模型构建和语料库开发等方式部分弥补其数据不足。尽管已有针对库尔德语的研究,但多集中于索拉尼(中央库尔德语)和库尔曼吉(北部库尔德语)两种方言。本文使用由两名母语者标注的6,854篇胡拉米语文本,涵盖15个类别,评估K近邻(KNN)、线性支持向量机(Linear SVM)、逻辑回归(LR)和决策树(DT)的文本分类表现。结果表明,线性SVM达到96%的准确率,显著优于其他方法,验证了该框架对濒危语言处理的有效性。
原文摘要 · Abstract (English)
Hawrami, a dialect of Kurdish, is classified as an endangered language as it suffers from the scarcity of data and the gradual loss of its speakers. Natural Language Processing projects can be used to partially compensate for data availability for endangered languages/dialects through a variety of approaches, such as machine translation, language model building, and corpora development. Similarly, NLP projects such as text classification are in language documentation. Several text classification studies have been conducted for Kurdish, but they were mainly dedicated to two particular dialects: Sorani (Central Kurdish) and Kurmanji (Northern Kurdish). In this paper, we introduce various text classification models using a dataset of 6,854 articles in Hawrami labeled into 15 categories by two native speakers. We use K-nearest Neighbor (KNN), Linear Support Vector Machine (Linear SVM), Logistic Regression (LR), and Decision Tree (DT) to evaluate how well those methods perform the classification task. The results indicate that the Linear SVM achieves a 96% of accuracy and outperforms the other approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。