用三角色自博弈强化学习提升大模型安全性,无需人工标注。
TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment
- 三角色闭环自博弈,自动优化攻击、防御与评估能力。
- 攻击者效果提升20%-50%,防御者安全性能增10%-30%。
- 适合关注大模型安全对齐与自动化训练的研究者。
近年来,大语言模型的安全风险日益突出,亟需减少有害内容生成。主流安全对齐方法通常采用三角色协同框架:攻击者生成对抗提示,防御者进行安全防护,评估者判断响应质量。本文提出一种闭环强化学习框架TriPlay-RL,实现三角色在近零人工标注下的持续协同进化。实验表明,攻击者在保持高输出多样性的同时,对抗有效性提升20%-50%;防御者在不损害通用推理能力的前提下,安全性能提升10%-30%;评估者通过迭代不断精进细粒度判别能力,可准确区分不安全回复、简单拒绝与有用引导。整体上,该框架建立了一种高效可扩展的大模型安全对齐范式,支持统一学习环内的持续共演化。
原文摘要 · Abstract (English)
In recent years, safety risks associated with large language models have become increasingly prominent, highlighting the urgent need to mitigate the generation of toxic and harmful content. The mainstream paradigm for LLM safety alignment typically adopts a collaborative framework involving three roles: an attacker for adversarial prompt generation, a defender for safety defense, and an evaluator for response assessment. In this paper, we propose a closed-loop reinforcement learning framework called TriPlay-RL that enables iterative and co-improving collaboration among three roles with near-zero manual annotation. Experimental results show that the attacker preserves high output diversity while achieving a 20%-50% improvement in adversarial effectiveness; the defender attains 10%-30% gains in safety performance without degrading general reasoning capability; and the evaluator continuously refines its fine-grained judgment ability through iterations, accurately distinguishing unsafe responses, simple refusals, and useful guidance. Overall, our framework establishes an efficient and scalable paradigm for LLM safety alignment, enabling continuous co-evolution within a unified learning loop.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。