首个中文网络欺凌事件检测数据集,按真实事件构建。
Chinese Cyberbullying Detection: Dataset, Method, and Validation
- 按真实事件组织标注,突破传统语气极性分类
- 构建91个事件、22万条评论的CHNCI数据集
- 结合解释生成与人工校验,确保标注质量
现有网络欺凌检测基准多基于话语极性(如‘攻击性’与‘非攻击性’),实为仇恨言论检测。但现实中,网络欺凌常通过具体事件引发社会关注。为此,本文提出一种基于事件的新型标注方法,构建首个中文网络欺凌事件检测数据集CHNCI,包含91个事件中的220,676条评论。首先融合三种基于解释生成的检测方法生成伪标签,再由人工标注者进行判断;并提出用于验证是否构成网络欺凌事件的评估标准。实验表明,该数据集可作为网络欺凌检测与事件预测任务的基准。据我们所知,这是首个针对中文网络欺凌事件检测的研究。
原文摘要 · Abstract (English)
Existing cyberbullying detection benchmarks were organized by the polarity of speech, such as "offensive" and "non-offensive", which were essentially hate speech detection. However, in the real world, cyberbullying often attracted widespread social attention through incidents. To address this problem, we propose a novel annotation method to construct a cyberbullying dataset that organized by incidents. The constructed CHNCI is the first Chinese cyberbullying incident detection dataset, which consists of 220,676 comments in 91 incidents. Specifically, we first combine three cyberbullying detection methods based on explanations generation as an ensemble method to generate the pseudo labels, and then let human annotators judge these labels. Then we propose the evaluation criteria for validating whether it constitutes a cyberbullying incident. Experimental results demonstrate that the constructed dataset can be a benchmark for the tasks of cyberbullying detection and incident prediction. To the best of our knowledge, this is the first study for the Chinese cyberbullying incident detection task.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。