arXiv:2505.15297cs.CL2025-05EMNLP被引 4

中文辱骂语净化需保留情绪基调,该研究提出首个针对性数据集。

Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites

  • 构建中文情感极性一致的改写数据集,标注1556组毒言-非毒重写对
  • 17个大模型在表情符号和谐音等隐性毒性上净化准确率不足60%
  • 适合做中文社交平台内容安全、情感敏感型文本生成的研究者

在保持说话人原意的前提下净化攻击性语言,是提升网络互动质量的关键挑战。尽管大语言模型在改写有毒内容方面展现出潜力,但常过度采用礼貌化表达,扭曲情感基调与沟通意图。这一问题在中文中尤为突出,因毒性常通过表情符号、谐音或语境隐含产生。本文提出 ToxiRewriteCN,首个专为保持情感极性设计的中文净化数据集,包含1556个精心标注的三元组,每组含一句有毒语句、一个情感一致的非毒改写句及标注的毒性片段。涵盖五类真实场景:常规表达、表情符号引发的毒性、谐音毒性,以及单轮与多轮对话。我们在四个维度(净化准确率、流畅性、内容保留度、情感极性)评估了17种大模型,包括商业与开源模型。结果显示,尽管商业模型和MoE架构表现最优,但在表情符号、谐音及对话类输入等更隐蔽或依赖上下文的情境中,所有模型均难以兼顾安全性与情感真实性。本研究公开 ToxiRewriteCN 数据集,以推动面向中文的可控、情感感知型净化技术发展。

原文摘要 · Abstract (English)

Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent. This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context. We present ToxiRewriteCN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity. The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans. It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues. We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity. Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs. We release ToxiRewriteCN to support future research on controllable, sentiment-aware detoxification for Chinese.

中文净化情感极性大模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。