首个面向混合语种泰卢固语的暴力语言识别数据集及模型研究
Overcoming Low-Resource Barriers in Tulu: Neural Models and Corpus Creation for OffensiveLanguage Identification
- 构建首个泰卢固语混合语社交文本标注数据集,涵盖四类标签
- 双向GRU自注意力模型达82%准确率,远超主流Transformer模型
- 为低资源、混合语种语言提供可复用的标注与建模框架
泰卢固语是印度南部主要使用的低资源德拉维达语系语言,尽管其数字存在感日益增强,但计算资源仍十分有限。本研究首次为代码混合的泰卢固语社交媒体内容构建了暴力语言识别(OLI)基准数据集,数据源自跨领域的YouTube评论。该数据集经高一致性标注(Krippendorff's alpha = 0.984),包含3,845条评论,分为四类:非暴力、非泰卢固语、无目标暴力、有目标暴力。我们评估了多种深度学习模型,包括GRU、LSTM、BiGRU、BiLSTM、CNN以及基于注意力机制的变体,还对比了mBERT和XLM-RoBERTa等变压器架构。其中,带自注意力的BiGRU模型表现最佳,达到82%准确率和0.81宏平均F1得分。相比之下,变压器模型表现不佳,凸显多语言预训练在代码混合、低资源语境下的局限性。该工作为泰卢固语及其他类似语言的NLP研究奠定了基础。
原文摘要 · Abstract (English)
Tulu, a low-resource Dravidian language predominantly spoken in southern India, has limited computational resources despite its growing digital presence. This study presents the first benchmark dataset for Offensive Language Identification (OLI) in code-mixed Tulu social media content, collected from YouTube comments across various domains. The dataset, annotated with high inter-annotator agreement (Krippendorff's alpha = 0.984), includes 3,845 comments categorized into four classes: Not Offensive, Not Tulu, Offensive Untargeted, and Offensive Targeted. We evaluate a suite of deep learning models, including GRU, LSTM, BiGRU, BiLSTM, CNN, and attention-based variants, alongside transformer architectures (mBERT, XLM-RoBERTa). The BiGRU model with self-attention achieves the best performance with 82% accuracy and a 0.81 macro F1-score. Transformer models underperform, highlighting the limitations of multilingual pretraining in code-mixed, under-resourced contexts. This work lays the foundation for further NLP research in Tulu and similar low-resource, code-mixed languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。