arXiv:2505.16722cs.CLcs.AI2025-05中稿 · COLM被引 3

跨语言去毒技术让低资源语言模型也能自动过滤毒性内容

Breaking mBad! Supervised Fine-tuning for Cross-Lingual Detoxification

  • 用监督微调实现高/低资源语言间毒性过滤能力迁移
  • 392组实验显示跨语言去毒在数据少时仍有效
  • 适合关注多语言安全与模型鲁棒性的研究者

随着大语言模型在各类全球应用中普及,确保其在不同语言环境下均无毒性仍是关键挑战。本文探索‘跨语言去毒’范式,通过跨语言迁移机制,在不同书写体系的高资源与低资源语言间实现毒性缓解能力传递。我们通过392组大规模实验评估了在数据有限条件下跨分布场景中的毒性降低效果,并分析了去毒对非毒性任务性能的影响,揭示了安全性与知识保留之间的权衡关系。代码与数据集已公开于 https://github.com/himanshubeniwal/Breaking-mBad。

原文摘要 · Abstract (English)

As large language models (LLMs) become increasingly prevalent in global applications, ensuring that they are toxicity-free across diverse linguistic contexts remains a critical challenge. We explore "Cross-lingual Detoxification", a cross-lingual paradigm that mitigates toxicity, enabling detoxification capabilities to transfer between high and low-resource languages across different script families. We analyze cross-lingual detoxification's effectiveness through 392 extensive settings to evaluate toxicity reduction in cross-distribution settings with limited data and investigate how mitigation impacts model performance on non-toxic tasks, revealing trade-offs between safety and knowledge preservation. Our code and dataset are publicly available at https://github.com/himanshubeniwal/Breaking-mBad.

跨语言去毒LLM安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。