arXiv:2506.10960cs.CLcs.AI2025-06被引 7

构建首个中文有害内容检测基准,提升模型识别能力。

ChineseHarm-Bench: A Chinese Harmful Content Detection Benchmark

  • 基于真实数据构建六类中文有害内容标注集。
  • 提出融合人工规则与大模型隐性知识的增强基线方法。
  • 适合中文内容安全研究者与平台审核系统开发者使用。

大型语言模型(LLMs)在自动化有害内容检测中应用日益广泛,有助于审核人员识别违规内容并提升内容审查的效率与准确性。然而,现有资源主要聚焦英文,中文数据集稀缺且范围有限。本文提出一个全面、专业标注的中文有害内容检测基准,涵盖六类代表性类别,数据全部来自真实场景。标注过程还生成了知识规则库,为大模型提供明确的专家知识支持。此外,我们提出一种知识增强基线方法,结合人工标注规则与大模型隐式知识,使小型模型性能达到领先大模型水平。代码与数据已开源:https://github.com/zjunlp/ChineseHarm-bench。

原文摘要 · Abstract (English)

Large language models (LLMs) have been increasingly applied to automated harmful content detection tasks, assisting moderators in identifying policy violations and improving the overall efficiency and accuracy of content review. However, existing resources for harmful content detection are predominantly focused on English, with Chinese datasets remaining scarce and often limited in scope. We present a comprehensive, professionally annotated benchmark for Chinese content harm detection, which covers six representative categories and is constructed entirely from real-world data. Our annotation process further yields a knowledge rule base that provides explicit expert knowledge to assist LLMs in Chinese harmful content detection. In addition, we propose a knowledge-augmented baseline that integrates both human-annotated knowledge rules and implicit knowledge from large language models, enabling smaller models to achieve performance comparable to state-of-the-art LLMs. Code and data are available at https://github.com/zjunlp/ChineseHarm-bench.

内容安全中文NLP大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。