arXiv:2501.01830cs.CRcs.AI2025-01被引 10

用强化学习自动探索复杂攻击策略,高效发现大模型安全漏洞

Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models

  • 基于强化学习自动优化恶意查询策略,动态适应防御变化
  • 探索效率提升,漏洞检测成功率高出16.63%、速度更快
  • 适合安全研究人员和模型评测团队快速识别复杂漏洞

自动化红队测试已成为发现大语言模型(LLMs)漏洞的关键方法。然而,现有方法多聚焦于孤立的安全缺陷,难以应对动态防御并高效挖掘复杂漏洞。为此,我们提出Auto-RT,一个基于强化学习的框架,可自动探索并优化复杂攻击策略,通过恶意查询有效暴露安全漏洞。具体设计两个关键机制:1)早期终止探索,聚焦高潜力攻击路径以加速搜索;2)带有中间降级模型的渐进式奖励追踪算法,动态调整搜索方向以实现漏洞成功利用。在多种LLM上的大量实验表明,Auto-RT显著提升了探索效率,自动优化攻击策略,检测范围更广,相比现有方法检测速度更快,成功率提高16.63%。

原文摘要 · Abstract (English)

Automated red-teaming has become a crucial approach for uncovering vulnerabilities in large language models (LLMs). However, most existing methods focus on isolated safety flaws, limiting their ability to adapt to dynamic defenses and uncover complex vulnerabilities efficiently. To address this challenge, we propose Auto-RT, a reinforcement learning framework that automatically explores and optimizes complex attack strategies to effectively uncover security vulnerabilities through malicious queries. Specifically, we introduce two key mechanisms to reduce exploration complexity and improve strategy optimization: 1) Early-terminated Exploration, which accelerate exploration by focusing on high-potential attack strategies; and 2) Progressive Reward Tracking algorithm with intermediate downgrade models, which dynamically refine the search trajectory toward successful vulnerability exploitation. Extensive experiments across diverse LLMs demonstrate that, by significantly improving exploration efficiency and automatically optimizing attack strategies, Auto-RT detects a boarder range of vulnerabilities, achieving a faster detection speed and 16.63\% higher success rates compared to existing methods.

模型安全红队测试强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。