arXiv:2608.12361cs.CLcs.CY2026-08ACL

用网络搜索增强大模型,识别中文新词中的隐性恶意用法

New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs

论文配图:New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
图 1 · 摘自论文原文
  • 基于公众共识构建毒害性新词分类体系与词库
  • 引入搜索增强框架SeTox,让小模型也能实时获取网络语境
  • 适用于内容审核系统,尤其适合追踪新兴网络污名化用语

新词(neologisms)在形式或含义上的新出现,可能成为新的毒害表达载体,如“乡村女孩”被用来贬低女性主义。这类新词表面中性,但在公共共识中已演变为有毒用法,对内容审核系统构成挑战且研究不足。本文提出一种分类体系,涵盖毒害性新词的来源及共识验证标准,并构建覆盖广泛风险类别的词库。为捕捉基于公众共识的毒性,提出SeTox框架——通过引入实时网络搜索,使静态大语言模型具备动态上下文感知能力。实验表明,即使使用30亿参数模型,SeTox仍优于近期大规模模型,证明其在融入现实知识方面的可扩展性。注意:本文包含可能令人不适的敏感内容。

原文摘要 · Abstract (English)

Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "country girl" as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in public consensus, posing challenges to moderation systems and remaining underexplored. In this paper, we investigate how to detect implicit toxicity expressed via neologisms. We first propose a taxonomy that captures the origins and consensus-verification criteria of toxic neologisms, followed by the construction of a lexicon spanning widely observed risk categories. To capture toxicity grounded in public consensus, we introduce SeTox, a search-augmented framework that enables static large language models (LLMs) to incorporate real-time web context for neologism toxicity detection. Experiments show that SeTox, even with 3B-scale models, outperforms recent large-scale models, demonstrating its scalability to incorporate real-world knowledge for toxic neologism detection. Disclaimer: this paper has offensive contents that may be disturbing to some readers.

新词检测毒性识别搜索增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。