arXiv:2508.05775cs.CLcs.CY2025-08综述被引 11

系统梳理大模型生成有害内容的威胁与防护策略

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM

  • 构建统一分类体系,涵盖无意毒性与恶意越狱攻击
  • 分析检测、分类、防御等多类安全机制的有效性
  • 适合关注AI伦理与安全的开发者及研究人员

大型语言模型(LLMs)在数字平台内容生成中带来革命性突破,同时可能无意生成有毒、冒犯或偏见内容。这种双重角色——既是强大文本生成工具,又是潜在有害语言源头——构成紧迫的社会技术挑战。本文系统综述近期研究,涵盖无意毒性、对抗性越狱攻击及全面缓解策略。探讨了LLM作为危害生成者与安全守护者的双重身份,通过检测、分类、内容审核和预防实现安全增强。提出一个统一的LLM相关危害与防御分类体系,分析新兴多模态及基于LLM的越狱策略,评估包括基于人类反馈的强化学习(RLHF)、提示工程与安全对齐在内的缓解方法。指出当前评估方法的局限性,并展望未来研究方向,以推动稳健且符合伦理的语言技术发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have revolutionized content creation across digital platforms, offering unprecedented capabilities in natural language generation and understanding. Meanwhile, they pose risks by inadvertently producing toxic, offensive, or biased content. This dual role of LLMs, both as powerful tools for text generation and as potential sources of harmful language, presents a pressing sociotechnical challenge. In this survey, we systematically review recent studies encompassing unintentional toxicity, adversarial jailbreak attacks, and comprehensive mitigation strategies. We explore LLMs' dual role as both generators of harm and enablers of safety through detection, classification, content moderation, and prevention. We propose a unified taxonomy of LLM-related harms and defenses, analyze emerging multimodal and LLM-assisted jailbreak strategies, and assess mitigation efforts, including reinforcement learning with human feedback (RLHF), prompt engineering, and safety alignment. Our review highlights the evolving landscape of LLM safety and identifies limitations in current evaluation methodologies. Ultimately, our review outlines future research directions to guide the development of robust and ethically aligned language technologies.

大模型安全有害内容越狱攻击伦理对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。