用微调Transformer模型提升低资源孟加拉语仇恨言论识别效果
Bangla Hate Speech Classification with Fine-tuned Transformer Models
- 采用多种Transformer模型对比,重点测试孟加拉语专用预训练模型
- 孟加拉语BERT在两项任务中表现最佳,优于m-BERT和XLM-RoBERTa
- 证明语言特化预训练对低资源语言至关重要,适合多语言NLP研究者
低资源语言的仇恨言论识别因数据不足、拼写多样性和语言差异而困难重重。孟加拉语是孟加拉国和印度西孟加拉邦超过2.3亿人的母语,尽管社交媒体自动化内容审核需求日益增长,但该语言在计算资源中仍严重缺失。本文研究了BLP 2025共享任务中的子任务1A和1B。我们复现了官方基线(如多数类、随机、支持向量机),并引入逻辑回归、随机森林和决策树作为基线方法。同时使用DistilBERT、BanglaBERT、m-BERT和XLM-RoBERTa等Transformer模型进行分类。所有Transformer模型在两个子任务中均优于传统基线,除DistilBERT外。其中,孟加拉语BERT在两项任务中表现最佳。尽管规模更小,其性能仍优于m-BERT和XLM-RoBERTa,表明语言特化预训练极为重要。结果凸显了为低资源孟加拉语开发预训练语言模型的潜力与必要性。
原文摘要 · Abstract (English)
Hate speech recognition in low-resource languages remains a difficult problem due to insufficient datasets, orthographic heterogeneity, and linguistic variety. Bangla is spoken by more than 230 million people of Bangladesh and India (West Bengal). Despite the growing need for automated moderation on social media platforms, Bangla is significantly under-represented in computational resources. In this work, we study Subtask 1A and Subtask 1B of the BLP 2025 Shared Task on hate speech detection. We reproduce the official baselines (e.g., Majority, Random, Support Vector Machine) and also produce and consider Logistic Regression, Random Forest, and Decision Tree as baseline methods. We also utilized transformer-based models such as DistilBERT, BanglaBERT, m-BERT, and XLM-RoBERTa for hate speech classification. All the transformer-based models outperformed baseline methods for the subtasks, except for DistilBERT. Among the transformer-based models, BanglaBERT produces the best performance for both subtasks. Despite being smaller in size, BanglaBERT outperforms both m-BERT and XLM-RoBERTa, which suggests language-specific pre-training is very important. Our results highlight the potential and need for pre-trained language models for the low-resource Bangla language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。