预测仇恨言论受害者对反驳言论的反应,揭示其是否回怼及回怼内容是否仍具攻击性。
Echoes of Discord: Forecasting Hater Reactions to Counterspeech
- 从仇恨者视角分析反驳言论影响,构建三轮对话数据集
- 三分类模型优于分步预测法,准确识别仇恨者是否回帖及回帖性质
- 发现反驳话语的语言特征可影响仇恨者反应,适合社交平台安全研究者
仇恨言论破坏在线社区的包容性并传播负面情绪与分裂。反驳言论被视为缓解其危害的重要手段。尽管已有研究探讨用户生成的反驳对社交媒体的影响,但极少关注仇恨者对反驳言论的具体反应,而这种即时态度转变正是反驳有效性的重要体现。本研究填补该空白,从仇恨者视角出发,分析反驳言论是否促使仇恨者重新参与对话,以及其回帖是否仍带有仇恨性质。我们构建了Reddit仇恨回响数据集(ReEco),包含三轮对话样本,用于评估反驳效果。为预测仇恨者行为,采用两种策略:两阶段反应预测器和三分类器。语言学分析揭示了不同反驳话语引发仇恨者不同反应的语言特征。实验表明,三分类模型优于先预测回帖再分类的两阶段方法。最后,我们总结了最佳模型的常见错误类型。
原文摘要 · Abstract (English)
Hate speech (HS) erodes the inclusiveness of online users and propagates negativity and division. Counterspeech has been recognized as a way to mitigate the harmful consequences. While some research has investigated the impact of user-generated counterspeech on social media platforms, few have examined and modeled haters' reactions toward counterspeech, despite the immediate alteration of haters' attitudes being an important aspect of counterspeech. This study fills the gap by analyzing the impact of counterspeech from the hater's perspective, focusing on whether the counterspeech leads the hater to reenter the conversation and if the reentry is hateful. We compile the Reddit Echoes of Hate dataset (ReEco), which consists of triple-turn conversations featuring haters' reactions, to assess the impact of counterspeech. To predict haters' behaviors, we employ two strategies: a two-stage reaction predictor and a three-way classifier. The linguistic analysis sheds insights on the language of counterspeech to hate eliciting different haters' reactions. Experimental results demonstrate that the 3-way classification model outperforms the two-stage reaction predictor, which first predicts reentry and then determines the reentry type. We conclude the study with an assessment showing the most common errors identified by the best-performing model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。