构建六维评估体系,量化分析反仇恨言论的有效性
Effectiveness of Counter-Speech against Abusive Content: A Multidimensional Annotation and Classification Study
- 从语言学与传播学出发,定义六维评价标准
- 在4214条数据上实现0.94~0.96的高准确率分类
- 为平台治理提供可操作的反仇恨言论评估工具
反仇恨言论(Counter-speech, CS)是缓解网络仇恨言论(HS)的关键策略,但其有效性评估标准仍不明确。本文提出一种基于语言学、传播学与论证理论的计算框架,定义了清晰性、证据支持、情感诉求、反驳力度、受众适配与公平性六个核心维度。基于此框架,对来自两个基准数据集的4,214条CS实例进行标注,形成公开可用的语料资源。同时提出多任务与依赖关系驱动的分类策略,在专家与用户撰写的CS中均取得0.94与0.96的平均F1值,显著优于基线模型,并揭示各维度间存在强相关性。
原文摘要 · Abstract (English)
Counter-speech (CS) is a key strategy for mitigating online Hate Speech (HS), yet defining the criteria to assess its effectiveness remains an open challenge. We propose a novel computational framework for CS effectiveness classification, grounded in linguistics, communication and argumentation concepts. Our framework defines six core dimensions - Clarity, Evidence, Emotional Appeal, Rebuttal, Audience Adaptation, and Fairness - which we use to annotate 4,214 CS instances from two benchmark datasets, resulting in a novel linguistic resource released to the community. In addition, we propose two classification strategies, multi-task and dependency-based, achieving strong results (0.94 and 0.96 average F1 respectively on both expert- and user-written CS), outperforming standard baselines, and revealing strong interdependence among dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。