arXiv:2507.07640cs.CL2025-07EMNLP被引 8

破解拼音伪装的中文辱骂内容,提升真实场景下的检测能力

Lost in Pronunciation: Detecting Chinese Offensive Language Disguised by Phonetic Cloaking Replacement

  • 构建四类拼音伪装文本分类体系,基于真实社交平台数据
  • 现有大模型在真实数据上F1仅0.672,零样本提示反而更差
  • 重用拼音提示策略显著提升检测效果,适合内容安全研究者

拼音伪装替换(Phonetic Cloaking Replacement, PCR)指用户故意使用同音或近音词来隐藏攻击性意图,已成为中文内容审核的主要障碍。尽管该问题广受关注,现有评估多依赖规则生成的合成扰动,忽视真实用户的创造力。本文提出四类表面形式的PCR分类体系,并构建了包含500条真实来源、经拼音伪装的辱骂帖子的数据集 extsc{ours},数据来自红笔记平台。在该数据集上对主流大模型进行基准测试发现,表现最好的模型F1仅为0.672,而零样本链式思考提示甚至导致性能下降。通过错误分析,我们重新评估此前被认为无效的拼音提示策略,结果表明其能有效恢复大量准确率。本研究首次提供中文PCR的完整分类体系,建立真实场景下的评测基准,并提出一种轻量级缓解方法,推动鲁棒毒性检测研究进展。

原文摘要 · Abstract (English)

Phonetic Cloaking Replacement (PCR), defined as the deliberate use of homophonic or near-homophonic variants to hide toxic intent, has become a major obstacle to Chinese content moderation. While this problem is well-recognized, existing evaluations predominantly rely on rule-based, synthetic perturbations that ignore the creativity of real users. We organize PCR into a four-way surface-form taxonomy and compile \ours, a dataset of 500 naturally occurring, phonetically cloaked offensive posts gathered from the RedNote platform. Benchmarking state-of-the-art LLMs on this dataset exposes a serious weakness: the best model reaches only an F1-score of 0.672, and zero-shot chain-of-thought prompting pushes performance even lower. Guided by error analysis, we revisit a Pinyin-based prompting strategy that earlier studies judged ineffective and show that it recovers much of the lost accuracy. This study offers the first comprehensive taxonomy of Chinese PCR, a realistic benchmark that reveals current detectors' limits, and a lightweight mitigation technique that advances research on robust toxicity detection.

中文NLP内容安全拼音伪装毒性检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。