中文有毒内容检测面临多重扰动挑战,模型易误判。
Exploring Multimodal Challenges in Toxic Chinese Detection: Taxonomy, Benchmark, and Findings
- 构建中文毒内容多模态扰动分类体系
- 9个主流大模型在扰动文本上检测率下降超30%
- 小样本微调反而引发大量误报,需谨慎使用
使用语言模型检测有毒内容虽重要但具挑战性。尽管大语言模型(LLMs)在理解中文方面表现强劲,但近期研究表明,对中文有毒文本进行简单字符替换即可轻易误导当前最先进的LLMs。本文强调中文语言的多模态特性是部署LLMs于中文有毒内容检测中的关键挑战。首先,我们提出一个包含3种扰动策略和8种具体方法的分类体系;随后,基于该分类体系构建数据集,并对9个来自中美两国的SOTA LLMs进行基准测试,评估其对被扰动的中文有毒文本的检测能力。此外,我们探索了上下文学习(ICL)和监督微调(SFT)等低成本增强方案。结果揭示两个关键发现:(1) LLMs对被扰动的多模态中文有毒内容检测能力显著下降;(2) 使用少量扰动样本进行ICL或SFT可能导致模型“过度纠正”,将大量正常中文内容错误识别为有毒。
原文摘要 · Abstract (English)
Detecting toxic content using language models is important but challenging. While large language models (LLMs) have demonstrated strong performance in understanding Chinese, recent studies show that simple character substitutions in toxic Chinese text can easily confuse the state-of-the-art (SOTA) LLMs. In this paper, we highlight the multimodal nature of Chinese language as a key challenge for deploying LLMs in toxic Chinese detection. First, we propose a taxonomy of 3 perturbation strategies and 8 specific approaches in toxic Chinese content. Then, we curate a dataset based on this taxonomy, and benchmark 9 SOTA LLMs (from both the US and China) to assess if they can detect perturbed toxic Chinese text. Additionally, we explore cost-effective enhancement solutions like in-context learning (ICL) and supervised fine-tuning (SFT). Our results reveal two important findings. (1) LLMs are less capable of detecting perturbed multimodal Chinese toxic contents. (2) ICL or SFT with a small number of perturbed examples may cause the LLMs "overcorrect'': misidentify many normal Chinese contents as toxic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。