arXiv:2607.11597cs.CL2026-07

Bangla仇恨言论检测模型在隐含表达上严重失效,尤其对讽刺和表情符号敏感。

Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection

  • 用多源数据训练模型,测试其在真实社交平台上的泛化能力。
  • 模型对隐含仇恨言论的识别率从91%降至63%,表情符号提升检测效果12%。
  • 研究揭示文化语境与表情符号对低资源语言检测的关键影响,适合政策与平台参考。

社交媒体平台上的仇恨言论(HS)传播威胁在线安全与伦理监管。自动检测在资源匮乏的语言如孟加拉语中尤为困难,因涉及文化背景、隐含表达与非正式语言模式。本研究通过诊断基准模型在隐含、依赖上下文的仇恨言论上失败原因,揭示孟加拉语仇恨言论检测的危机。六种架构(FastText + CNN, FastText + LSTM, FastText + BiLSTM, BanglaBERT, BanglaBERT + CNN, BanglaBERT + BiLSTM)在约7.5万条基准数据与约12万条多源数据上训练,外部分别在约200条来自Facebook、Twitter、YouTube的标注数据集上验证,其中仇恨言论分为显性和隐性。BanglaBERT在基准数据集上F1为91.4%,但在外部数据集下降至75.3%,对含讽刺与表情符号的隐性仇恨言论仅为63.4%。FastText + CNN准确率从78.0%降至51.2%。引入表情符号感知预处理可使隐性仇恨言论检测提升最高12%,而移除表情符号导致性能显著下降(F1:0.75→0.63)。政治性或讽刺性评论常被误判,暴露过度监管风险。研究不仅揭示了由隐含、文化嵌入与表情符号构成的泛化危机,更强调需构建适应性强、表情符号感知且文化契合的框架,以实现伦理监管与言论自由的平衡。研究成果为研究人员、社交平台及政策制定者提供面向低资源语言的上下文敏感检测系统设计思路。

原文摘要 · Abstract (English)

The spread of hate speech (HS) across different social media platforms (SMPs) poses a major concern for online safety and ethical moderation. Automatic detection of HS remains a challenging task, especially in under-resourced languages like Bangla, due to cultural context, implicit expressions, and informal linguistic patterns. This study aimed to expose the crisis of Bangla HS detection systems by diagnosing how and why benchmark-trained models fail to identify implicit, context-dependent HS. Six architectures (FastText + CNN, FastText + LSTM, FastText + BiLSTM, BanglaBERT, BanglaBERT + CNN, and BanglaBERT + BiLSTM) were trained on benchmark datasets (about 75,000 posts) and a merged multi-source dataset (about 120,000 posts), then externally validated on an annotated dataset (about 200 posts) collected from Facebook, Twitter, and YouTube, labeled as HS and non-HS, where HS was further categorized as explicit and implicit. BanglaBERT achieved an F1-score of 91.4% on benchmark datasets but declined to 75.3% on the external set and 63.4% for implicit HS involving sarcasm and emojis. The accuracy of FastText + CNN dropped from 78.0% to 51.2% under similar conditions. Emoji-aware preprocessing improved implicit HS detection by up to 12%, whereas emoji removal caused a notable decline in performance (F1: 0.75 to 0.63). Frequent misclassifications in politically charged or satirical comments revealed over-policing risks. This study not only exposes the generalization crisis due to implicit, culturally embedded, and emoji-laden expressions but also underscores the need for developing adaptive, emoji-aware, and culturally grounded frameworks that ensure ethical moderation while preserving freedom of expression. Findings of this study provide insights for researchers, SMPs, and policymakers to design more context-sensitive HS detection systems for low-resource languages.

仇恨言论低资源语言表情符号文化语境

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。