arXiv:2606.25380cs.CL2026-06中稿 · ACL综述

梳理多语言大模型毒性检测与治理策略,揭示跨语言安全短板

A Survey of Toxicity Detection and Mitigation Strategies for Multilingual Language Models

  • 系统整理多语言场景下毒性攻击的七种威胁模式
  • 发现当前方法在低资源语言覆盖上存在明显不足
  • 适合关注AI安全、多语言伦理的研究者与从业者

大语言模型正日益跨语言部署,但其安全性在不同语言和文化背景下表现不一。本综述整合了多语言大模型毒性检测与净化的研究工作。首先梳理了利用语言选择、翻译中转、语码转换、拼写变异、多轮交互及部署后微调等手段削弱安全对齐的威胁模型。随后归纳了三类任务形式(毒性到中性重写、毒性分类、生成毒性评估),四种多语言检测方法(跨语言编码器、翻译流水线、表征探测、基于大模型的检测器),以及数据过滤、监督与偏好微调、解码时引导、表征编辑、多语言防护墙等缓解策略。在这些领域中,我们识别出持续挑战:语言覆盖不均、危害定义受文化影响、评估标准碎片化,以及净化过程可能抑制合法方言或身份表达的风险。

原文摘要 · Abstract (English)

Large language models (LLMs) are increasingly deployed across languages, but their safety behavior remains uneven across linguistic and cultural contexts. This survey synthesizes work on toxicity detection and detoxification for multilingual LLMs. We first catalogue threat models that exploit language choice, translation pivots, code-switching, orthographic variation, multi-turn interaction, and post-deployment fine-tuning to weaken safety alignment. We then organize task formulations (toxic-to-neutral rewriting, toxicity classification, and toxic-generation evaluation), multilingual detection approaches (cross-lingual encoders, translation pipelines, representation-level probes, and LLM-based detectors), and mitigation strategies spanning data filtering, supervised and preference-based tuning, decoding-time steering, representation editing, and multilingual guardrails. Across these areas, we identify persistent challenges: uneven language coverage, culturally contingent definitions of harm, fragmented evaluation protocols, and the risk that detoxification suppresses legitimate dialectal or identity-related expression.

AI安全多语言毒性检测大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。