arXiv:2502.10632cs.CL2025-02被引 2

用翻译+微调提升泰卢固语仇恨言论检测效果

Code-Mixed Telugu-English Hate Speech Detection

  • 将泰卢固语文本转为英语后用Transformer模型微调
  • 译后训练使德伯塔和印度仇恨-穆里尔模型表现最佳
  • 适合低资源语言的仇恨内容识别研究者参考

针对低资源语言泰卢固语的仇恨言论检测是自然语言处理中的重要挑战。本研究评估了包括TeluguHateBERT、HateBERT、DeBERTa、Muril、IndicBERT、Roberta和Hindi-Abusive-MuRIL在内的多种基于Transformer的模型在泰卢固语仇恨言论分类中的表现。通过使用低秩适配(LoRA)对模型进行微调以提升效率与性能,并探索将泰卢固语文本通过Google Translate转化为英语的多语言方法对分类准确率的影响。实验表明,多数模型在翻译后性能提升,其中DeBERTa和Hindi-Abusive-MuRIL在原始泰卢固语数据集及翻译数据集上均取得最高准确率与F1分数。值得注意的是,Hindi-Abusive-MuRIL在两种设置下均优于其他模型,展现出跨语言环境下的鲁棒性。结果说明,翻译可使模型利用英语中更丰富的语言特征,从而提高分类性能。这表明多语言处理是低资源语言仇恨言论检测的有效策略。

原文摘要 · Abstract (English)

Hate speech detection in low-resource languages like Telugu is a growing challenge in NLP. This study investigates transformer-based models, including TeluguHateBERT, HateBERT, DeBERTa, Muril, IndicBERT, Roberta, and Hindi-Abusive-MuRIL, for classifying hate speech in Telugu. We fine-tune these models using Low-Rank Adaptation (LoRA) to optimize efficiency and performance. Additionally, we explore a multilingual approach by translating Telugu text into English using Google Translate to assess its impact on classification accuracy. Our experiments reveal that most models show improved performance after translation, with DeBERTa and Hindi-Abusive-MuRIL achieving higher accuracy and F1 scores compared to training directly on Telugu text. Notably, Hindi-Abusive-MuRIL outperforms all other models in both the original Telugu dataset and the translated dataset, demonstrating its robustness across different linguistic settings. This suggests that translation enables models to leverage richer linguistic features available in English, leading to improved classification performance. The results indicate that multilingual processing can be an effective approach for hate speech detection in low-resource languages. These findings demonstrate that transformer models, when fine-tuned appropriately, can significantly improve hate speech detection in Telugu, paving the way for more robust multilingual NLP applications.

仇恨检测低资源语言多语言Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。