首个面向中文仇恨言论的细粒度标注数据集,助力精准识别攻击目标。
STATE ToxiCN: A Benchmark for Span-level Target-Aware Toxicity Extraction in Chinese Hate Speech Detection
- 构建首个中文仇恨言论跨度级标注数据集,含目标-论点-仇恨群体四元组
- 首次系统评估大模型对中文仇恨俚语的检测能力,发现现有模型表现有限
- 为中文仇恨言论研究提供可复用数据与评测基准,适合安全与自然语言处理研究者
仇恨言论的泛滥对社会造成严重危害,其强度与指向性紧密关联于具体目标与论述。然而,中文仇恨言论检测研究相对滞后,现有数据集缺乏细粒度的跨度级标注,且对中文仇恨俚语的研究不足。本文提出首个中文仇恨言论跨度级标注数据集——STATE ToxiCN,包含目标-论点-仇恨群体三元组(实际为四元组)标注。我们基于该数据集评估了现有模型在细粒度检测任务上的表现,并首次开展中文仇恨俚语研究,评估大语言模型对此类表达的识别能力。本工作为中文仇恨言论的细粒度检测提供了重要资源与洞见。
原文摘要 · Abstract (English)
The proliferation of hate speech has caused significant harm to society. The intensity and directionality of hate are closely tied to the target and argument it is associated with. However, research on hate speech detection in Chinese has lagged behind, and existing datasets lack span-level fine-grained annotations. Furthermore, the lack of research on Chinese hateful slang poses a significant challenge. In this paper, we provide a solution for fine-grained detection of Chinese hate speech. First, we construct a dataset containing Target-Argument-Hateful-Group quadruples (STATE ToxiCN), which is the first span-level Chinese hate speech dataset. Secondly, we evaluate the span-level hate speech detection performance of existing models using STATE ToxiCN. Finally, we conduct the first study on Chinese hateful slang and evaluate the ability of LLMs to detect such expressions. Our work contributes valuable resources and insights to advance span-level hate speech detection in Chinese.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。