arXiv:2507.16183cs.CL2025-07被引 11

首个多方言孟加拉语仇恨言论数据集,助力本土化内容审核

BIDWESH: A Bangla Regional Based Hate Speech Detection Dataset

  • 基于BD-SHS语料库,将9183条数据转译并标注三种主要方言
  • 每条数据标注仇恨存在性、类型及目标群体,支持细粒度分析
  • 填补低资源语言方言仇恨内容检测空白,适合本地化NLP研究

数字平台上的仇恨言论已成为全球关注问题,尤其在语言多样性突出的国家如孟加拉国,地方方言在日常交流中扮演重要角色。尽管标准孟加拉语的仇恨言论检测已有进展,现有数据集和系统仍无法覆盖巴里沙尔、诺阿哈尔、吉大港等方言中的非正式且富含文化特色的表达,导致检测能力有限且存在偏见,大量有害内容未被识别。为此,本研究提出BIDWESH——首个多方言孟加拉语仇恨言论数据集,通过将9,183条来自BD-SHS语料库的内容翻译并标注至三种主要区域方言,每条数据均经人工验证,标记仇恨是否存在、类型(诽谤、性别、宗教、煽动暴力)及目标群体(个人、男性、女性、群体),确保语言与语境准确性。该数据集为孟加拉语仇恨言论检测提供了语言丰富、平衡且包容的资源,推动方言敏感型自然语言处理工具的发展,显著促进低资源语言环境下的公平、情境感知的内容审核。

原文摘要 · Abstract (English)

Hate speech on digital platforms has become a growing concern globally, especially in linguistically diverse countries like Bangladesh, where regional dialects play a major role in everyday communication. Despite progress in hate speech detection for standard Bangla, Existing datasets and systems fail to address the informal and culturally rich expressions found in dialects such as Barishal, Noakhali, and Chittagong. This oversight results in limited detection capability and biased moderation, leaving large sections of harmful content unaccounted for. To address this gap, this study introduces BIDWESH, the first multi-dialectal Bangla hate speech dataset, constructed by translating and annotating 9,183 instances from the BD-SHS corpus into three major regional dialects. Each entry was manually verified and labeled for hate presence, type (slander, gender, religion, call to violence), and target group (individual, male, female, group), ensuring linguistic and contextual accuracy. The resulting dataset provides a linguistically rich, balanced, and inclusive resource for advancing hate speech detection in Bangla. BIDWESH lays the groundwork for the development of dialect-sensitive NLP tools and contributes significantly to equitable and context-aware content moderation in low-resource language settings.

仇恨言论孟加拉语方言数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。