用字符n-gram方法实现罗马尼亚语作者归属,效果媲美复杂模型。
Oldies but Goldies: The Potential of Character N-grams for Romanian Texts
- 用字符n-gram特征配合神经网络进行作者识别
- 5-gram特征下有四次实验达到100%准确率
- 适合资源少的低频语言研究者参考
本研究针对罗马尼亚语文本的作者归属问题,使用ROST语料库这一领域标准基准,系统评估了六种机器学习方法:支持向量机(SVM)、逻辑回归(LR)、k近邻(k-NN)、决策树(DT)、随机森林(RF)和人工神经网络(ANN),均采用字符n-gram特征进行分类。其中,ANN模型表现最佳,在使用5-gram特征时,十五次运行中有四次实现完全分类。结果表明,轻量级、可解释的字符n-gram方法可在罗马尼亚语作者归属任务中达到顶尖精度,与更复杂的模型相当。研究凸显了简单风格特征在资源受限或研究不足的语言环境中的潜力。
原文摘要 · Abstract (English)
This study addresses the problem of authorship attribution for Romanian texts using the ROST corpus, a standard benchmark in the field. We systematically evaluate six machine learning techniques: Support Vector Machine (SVM), Logistic Regression (LR), k-Nearest Neighbors (k-NN), Decision Trees (DT), Random Forests (RF), and Artificial Neural Networks (ANN), employing character n-gram features for classification. Among these, the ANN model achieved the highest performance, including perfect classification in four out of fifteen runs when using 5-gram features. These results demonstrate that lightweight, interpretable character n-gram approaches can deliver state-of-the-art accuracy for Romanian authorship attribution, rivaling more complex methods. Our findings highlight the potential of simple stylometric features in resource, constrained or under-studied language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。