arXiv:2501.15042cs.CL2025-01被引 17

首个面向中文网络暴力检测的细粒度会话数据集,支持精准识别每条评论性质。

SCCD: A Session-based Dataset for Chinese Cyberbullying Detection

  • 构建677个微博会话样本,每条评论标注细粒度标签而非简单二分类
  • 首次提供中文网络暴力检测的细粒度会话级数据集,填补领域空白
  • 适合研究中文社交媒体内容安全、情感分析与多粒度文本分类的学者使用

网络暴力内容的泛滥对社会福祉构成日益严重的威胁。然而,中文语境下的网络暴力检测研究仍不充分,主要受限于缺乏全面且可靠的语料库。目前尚无专门针对中文网络暴力检测的数据集,且现有会话级数据集在评论层级普遍缺乏细粒度标注。为此,我们提出一个新型中文网络暴力检测数据集SCCD,包含来自微博平台的677个会话级样本,每个会话中的评论均被赋予细粒度标签,而非传统二分类标签。通过实证评估多种基线模型在SCCD上的表现,揭示了有效进行中文网络暴力检测所面临的挑战。

原文摘要 · Abstract (English)

The rampant spread of cyberbullying content poses a growing threat to societal well-being. However, research on cyberbullying detection in Chinese remains underdeveloped, primarily due to the lack of comprehensive and reliable datasets. Notably, no existing Chinese dataset is specifically tailored for cyberbullying detection. Moreover, while comments play a crucial role within sessions, current session-based datasets often lack detailed, fine-grained annotations at the comment level. To address these limitations, we present a novel Chinese cyber-bullying dataset, termed SCCD, which consists of 677 session-level samples sourced from a major social media platform Weibo. Moreover, each comment within the sessions is annotated with fine-grained labels rather than conventional binary class labels. Empirically, we evaluate the performance of various baseline methods on SCCD, highlighting the challenges for effective Chinese cyberbullying detection.

网络暴力中文数据集细粒度标注会话建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。