不用复杂评分模型,用对抗训练让大模型更安全
PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training
- 用自动红队攻击发现漏洞,结合对抗训练提升防御能力
- 在5个主流大模型上测试,安全性能优于传统方法,计算成本降61%
- 适合资源有限的团队快速部署高安全性大模型
大型语言模型在众多应用中展现出强大能力,但其安全风险威胁关键场景的可靠部署。现有安全对齐方法多依赖过程奖励模型(PRM)评估中间推理步骤,带来显著计算开销与可扩展性瓶颈。本文提出一种无需PRM的安全对齐框架,通过自动化红队攻击与对抗训练实现强安全保障,同时保持计算高效。该方法利用遗传算法优化、多智能体仿真和高级提示变异技术系统识别模型漏洞,并通过课程学习与自适应正则化机制进行针对性对抗训练。在五个前沿大模型上的全面实验表明,本方法在安全对齐性能上优于基于PRM的方法,同时降低61%的计算成本。框架还集成透明报告与持续审计机制,支持迭代改进与合规管理。本工作推动了高效大模型安全对齐的发展,为资源受限机构提供可及的强安全方案,并为应对不断演化的对抗威胁奠定可扩展基础。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse applications, yet they pose significant security risks that threaten their safe deployment in critical domains. Current security alignment methodologies predominantly rely on Process Reward Models (PRMs) to evaluate intermediate reasoning steps, introducing substantial computational overhead and scalability constraints. This paper presents a novel PRM-free security alignment framework that leverages automated red teaming and adversarial training to achieve robust security guarantees while maintaining computational efficiency. Our approach systematically identifies vulnerabilities through sophisticated attack strategies including genetic algorithm optimization, multi-agent simulation, and advanced prompt mutation techniques. The framework enhances model robustness via targeted adversarial training with curriculum learning and adaptive regularization mechanisms. Comprehensive experimental evaluation across five state-of-the-art LLMs demonstrates that our method achieves superior security alignment performance compared to PRM-based approaches while reducing computational costs by 61\%. The framework incorporates transparent reporting and continuous audit mechanisms that enable iterative security improvement and regulatory compliance. Our contributions advance the field of efficient LLM security alignment by democratizing access to robust security measures for resource-constrained organizations and providing a scalable foundation for addressing evolving adversarial threats.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。