arXiv:2509.18562cs.MMcs.AI2025-09

提升中文居高临下语言检测准确率,助力平台内容治理

CPCLDETECTOR: Knowledge Enhancement and Alignment Selection for Chinese Patronizing and Condescending Language Detection

  • 构建含10.3万条评论的新数据集PCLMMPLUS,弥补原数据缺失
  • 提出CPCLDetector模型,在新数据集上性能超越现有最优方法
  • 适合从事中文有害内容检测与平台治理的研究者和工程师

中文居高临下语言(CPCL)是一种隐性歧视性有毒言论,针对视频平台上的弱势群体。现有数据集缺乏用户评论,而评论是视频内容的直接反映,其缺失削弱了模型对视频语境的理解,导致部分CPCL视频无法被识别。为此,本研究重构了包含10.3万条评论条目的新数据集PCLMMPLUS,显著扩充数据规模。同时提出CPCLDetector模型,集成对齐选择与知识增强的评论内容模块。大量实验表明,该模型在PCLMM和PCLMMPLUS上均优于当前最优方法,能更准确识别CPCL视频,支持内容治理并保护弱势群体。代码与数据集已公开于https://github.com/jiaxunyang256/PCLD。

原文摘要 · Abstract (English)

Chinese Patronizing and Condescending Language (CPCL) is an implicitly discriminatory toxic speech targeting vulnerable groups on Chinese video platforms. The existing dataset lacks user comments, which are a direct reflection of video content. This undermines the model's understanding of video content and results in the failure to detect some CPLC videos. To make up for this loss, this research reconstructs a new dataset PCLMMPLUS that includes 103k comment entries and expands the dataset size. We also propose the CPCLDetector model with alignment selection and knowledge-enhanced comment content modules. Extensive experiments show the proposed CPCLDetector outperforms the SOTA on PCLMM and achieves higher performance on PCLMMPLUS . CPLC videos are detected more accurately, supporting content governance and protecting vulnerable groups. Code and dataset are available at https://github.com/jiaxunyang256/PCLD.

文本检测有害内容中文NLP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。