建模标注分歧可提升仇恨言论检测效果
Dealing with Annotator Disagreement in Hate Speech Classification
- 用多种聚合策略处理标注者分歧,避免丢弃有用数据
- 利用标注者评分强度提升分类准确率,达新纪录
- 适合关注公平性与鲁棒性的内容安全研究者
仇恨言论检测在社交媒体中至关重要,但因主观性强,标注者常出现分歧,尤其针对模糊或边缘内容。传统方法要么丢弃不一致样本,要么强制设定“黄金标准”,忽视了不确定性与多元观点的价值。本文系统研究标注分歧问题,评估了多数投票、序数策略(最小值、最大值、均值)等聚合方法在二分类、四分类和六分类任务中的表现,并引入标注者感知的仇恨言论强度评分,探索回归与混合建模范式。结果表明,过滤不一致样本会导致结果过度乐观;而标注者强度评分能提供互补信号,显著提升性能。最终,在土耳其语推文上达成新的最先进水平,证明合理建模标注分歧可成为构建更稳健系统的宝贵资源。
原文摘要 · Abstract (English)
Hate speech detection is a crucial task, especially on social media where harmful content can spread quickly. Collecting social media content (tweets etc.) to train machine learning models is easy, but detecting and categorizing hate speech can be difficult due to the inherently subjective nature. This subjectivity leads to frequent disagreement among annotators, particularly for subtle or borderline content. Traditional approaches either discard non-consensus samples or force a ''gold standard'' through expert adjudication, ignoring valuable information about uncertainty and diverse human perspectives. We examine the largely overlooked problem of annotator disagreement in hate speech classification and evaluate a range of aggregation methods, including majority voting, ordinal strategies (minimum, maximum, and mean), and analyze their impact across binary, 4-class, and 6-class classification tasks. In addition, we leverage annotators' perceived hate speech strength scores to explore regression-based and hybrid modeling approaches. Among others, we show that filtering non-consensus samples results in over-optimistic results and that the perceived strength provides a complementary signal that enhance classification performance. Finally, we establish new state-of-the-art results for hate speech detection in Turkish tweets, and demonstrate that annotator disagreement, when properly modeled, is a valuable resource for building more robust and reliable systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。