Metis通过自我进化策略优化,高效突破大模型安全防护。
Metis: Learning to Jailbreak LLMs via Self-Evolving Metacognitive Policy Optimization

- 构建对抗性元认知闭环,动态诊断防御逻辑并优化攻击策略。
- 平均攻击成功率89.2%,在前沿模型上仍达76%~78%,显著优于传统方法。
- 降低8.2倍以上令牌开销,适合研究模型安全漏洞与防御机制者。
红队测试对发现大语言模型(LLM)漏洞至关重要。尽管自动化方法提升了可扩展性,但现有方法多依赖静态启发式或随机搜索,对先进安全对齐措施缺乏鲁棒性。为此,我们提出Metis,将越狱攻击重构为对抗性部分可观测马尔可夫决策过程(POMDP)中的推理时策略优化。Metis采用自进化元认知循环,对目标防御逻辑进行因果诊断,并利用结构化反馈作为语义梯度优化策略,通过透明的推理轨迹提升可解释性。在10种不同模型上的广泛评估表明,Metis在对比方法中实现了最高的平均攻击成功率达89.2%,在强韧前沿模型上表现依然出色(如O1为76.0%,GPT-5-chat为78.0%),而传统基线在此类模型上性能显著下降。通过以定向优化替代冗余探索,Metis平均减少8.2倍、最高达11.4倍的令牌消耗。分析显示,当前防御在受控闭环推理轨迹下仍易受内部引导攻击,凸显了需在推理阶段动态推理安全性的下一代防御体系的迫切需求。
原文摘要 · Abstract (English)
Red teaming is critical for uncovering vulnerabilities in Large Language Models (LLMs). While automated methods have improved scalability, existing approaches often rely on static heuristics or stochastic search, rendering them brittle against advanced safety alignment. To address this, we introduce Metis, a framework that reformulates jailbreaking as inference-time policy optimization within an adversarial Partially Observable Markov Decision Process (POMDP). Metis employs a self-evolving metacognitive loop to perform causal diagnosis of a target's defense logic and leverages structured feedback as a semantic gradient to refine its policy, offering enhanced interpretability through transparent reasoning traces. Extensive evaluations across 10 diverse models demonstrate that Metis achieves the strongest average Attack Success Rate (ASR) among compared methods at 89.2%, maintaining high efficacy on resilient frontier models (e.g., 76.0% on O1 and 78.0% on GPT-5-chat) where traditional baselines exhibit substantial performance degradation. By replacing redundant exploration with directed optimization, Metis reduces token costs by an average of 8.2x and up to 11.4x. Our analysis reveals that current defenses remain vulnerable to internally-steered, closed-loop reasoning trajectories under the tested settings, highlighting a critical need for next-generation defenses capable of reasoning about safety dynamically during inference.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。