arXiv:2511.02356cs.CRcs.LG2025-11ACL被引 1

自动发现并进化越狱攻击策略,提升攻击多样性与适应性。

ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs

  • 构建闭环机制,自动生成并提炼攻击策略。
  • 在黑盒环境下成功率显著高于现有基线。
  • 动态分层策略库,高效管理成功与失败模式。

尽管经过大量安全对齐,大型语言模型(LLMs)仍易受越狱攻击。现有方法普遍缺乏从交互中持续学习和自我演进的能力,限制了攻击策略的多样性和适应性。为此,我们提出 ASTRA,一个能够自主发现、检索和演化攻击策略的自动化框架。ASTRA 采用闭环「攻击-评估-提炼-重用」机制,不仅能生成攻击提示,还能从每次交互中自动提炼可复用的策略。为系统化管理这些策略,我们引入动态三层策略库(有效、有潜力、无效),根据表现对策略进行分类。该分层记忆机制使框架能利用成功模式提升效率,同时通过避免已知失败优化探索空间。大量黑盒环境实验表明,ASTRA 显著优于现有基线。

原文摘要 · Abstract (English)

Despite extensive safety alignment, Large Language Models (LLMs) remain vulnerable to jailbreak attacks. However, existing methods generally lack the capability for continuous learning and self-evolution from interactions, limiting the diversity and adaptability of attack strategies. To address this, we propose ASTRA, an automated framework capable of autonomously discovering, retrieving, and evolving attack strategies. ASTRA operates on a closed-loop ``attack-evaluate-distill-reuse'' mechanism, which not only generates attack prompts but also automatically distills reusable strategies from every interaction. To systematically manage these strategies, we introduce a dynamic three-tier strategy library (Effective, Promising, and Ineffective) that categorizes strategies based on performance. This hierarchical memory mechanism enables the framework to enhance efficiency by leveraging successful patterns while optimizing the exploration space by avoiding known failures. Extensive experiments in a black-box setting demonstrate that ASTRA significantly outperforms existing baselines.

越狱攻击自动化框架策略演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。