arXiv:2511.07304cs.CL2025-11被引 2

用多模型集成与多任务学习提升孟加拉语仇恨言论识别效果

Retriv at BLP-2025 Task 1: A Transformer Ensemble and Multi-Task Learning Approach for Bangla Hate Speech Identification

  • 采用多种孟加拉语Transformer模型软投票集成
  • 在三个子任务中分别取得72.6%以上的微平均F1分数
  • 结果对低资源语言仇恨言论检测具有参考价值

本文针对孟加拉语仇恨言论识别这一社会意义重大但语言挑战性强的任务展开研究。作为IJCNLP-AACL 2025年BLP研讨会共享任务的一部分,我们团队'Retriv'参与了全部三个子任务:(1A)仇恨类型分类,(1B)目标群体识别,(1C)类型、严重程度和目标的联合检测。对于1A和1B任务,我们采用BanglaBERT、MuRIL、IndicBERTv2等Transformer模型的软投票集成;对于1C任务,训练三种多任务变体并通过加权投票集成预测结果。系统在1A和1B任务上分别获得72.75%和72.69%的微平均F1分数,在1C任务上取得72.62%的加权微平均F1分数。在共享任务排行榜上分别位列第9、第10和第7名。这些结果表明,变压器模型集成与加权多任务框架在低资源语境下推动孟加拉语仇恨言论检测具有潜力。我们已公开实验脚本供社区使用。

原文摘要 · Abstract (English)

This paper addresses the problem of Bangla hate speech identification, a socially impactful yet linguistically challenging task. As part of the "Bangla Multi-task Hate Speech Identification" shared task at the BLP Workshop, IJCNLP-AACL 2025, our team "Retriv" participated in all three subtasks: (1A) hate type classification, (1B) target group identification, and (1C) joint detection of type, severity, and target. For subtasks 1A and 1B, we employed a soft-voting ensemble of transformer models (BanglaBERT, MuRIL, IndicBERTv2). For subtask 1C, we trained three multitask variants and aggregated their predictions through a weighted voting ensemble. Our systems achieved micro-f1 scores of 72.75% (1A) and 72.69% (1B), and a weighted micro-f1 score of 72.62% (1C). On the shared task leaderboard, these corresponded to 9th, 10th, and 7th positions, respectively. These results highlight the promise of transformer ensembles and weighted multitask frameworks for advancing Bangla hate speech detection in low-resource contexts. We made experimental scripts publicly available for the community.

仇恨言论识别孟加拉语多任务学习模型集成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。