首个中文仇恨言论细粒度数据集,提升识别与解释能力
Fine-Grained Chinese Hate Speech Understanding: Span-Level Resources, Coded Term Lexicon, and Enhanced Detection Frameworks
- 构建首个中文细粒度仇恨言论标注数据集,支持精准定位攻击目标
- 首次系统研究中文隐晦仇恨词汇,发现大模型解释力存在明显短板
- 融合人工标注词典提升检测效果,适合安全与内容审核场景
仇恨言论的泛滥对社会造成严重伤害,其强度与指向性高度依赖具体目标和论点。近年来,大量基于机器学习的方法被用于自动检测网络平台上的仇恨言论。然而,中文仇恨言论检测研究仍显滞后,可解释性研究面临两大挑战:一是缺乏细粒度的跨度级标注数据集,限制了模型对仇恨语义的深层理解;二是对隐晦仇恨词汇的识别与解释研究不足,制约了模型在复杂现实场景中的可解释性。为此,本文贡献如下:(1)提出首个中文跨度级目标感知仇恨言论提取数据集——STATE ToxiCN,用于评估现有模型的仇恨语义理解能力;(2)首次系统研究中文隐晦仇恨术语及大语言模型对其语义的解读能力;(3)提出一种将标注词典融入模型的方法,显著提升仇恨言论检测性能。本工作为推进中文仇恨言论检测的可解释性研究提供了重要资源与洞见。
原文摘要 · Abstract (English)
The proliferation of hate speech has inflicted significant societal harm, with its intensity and directionality closely tied to specific targets and arguments. In recent years, numerous machine learning-based methods have been developed to detect hateful comments on online platforms automatically. However, research on Chinese hate speech detection lags behind, and interpretability studies face two major challenges: first, the scarcity of span-level fine-grained annotated datasets limits models' deep semantic understanding of hate speech; second, insufficient research on identifying and interpreting coded hate speech restricts model explainability in complex real-world scenarios. To address these, we make the following contributions: (1) We introduce the Span-level Target-Aware Toxicity Extraction dataset (STATE ToxiCN), the first span-level Chinese hate speech dataset, and evaluate the hate semantic understanding of existing models using it. (2) We conduct the first comprehensive study on Chinese coded hate terms, LLMs' ability to interpret hate semantics. (3) We propose a method to integrate an annotated lexicon into models, significantly enhancing hate speech detection performance. Our work provides valuable resources and insights to advance the interpretability of Chinese hate speech detection research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。