arXiv:2605.27110cs.CRcs.CL2026-05

通过引导模型自我推理边界,逐步诱使其泄露敏感内容。

BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

论文配图:BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning
图 1 · 摘自论文原文
  • 分三步诱导模型自我揭示安全边界,利用其一致性倾向
  • 在多个评测集上对顶级大模型攻击成功率显著超越基线
  • 适合研究模型安全漏洞或对抗攻防的学者参考

本文提出BAIT(Boundary-Aware Iterative Trap)框架,一种三步式越狱攻击方法,通过引导模型内部披露实现恶意目标。首先要求模型识别防护边界,再促使其细化该边界,最后请求具体示例。每一步均基于前一响应迭代扩展,将模型自身的推理与一致性倾向转化为披露路径。在AdvBench、JailbreakBench、AIR-Bench和SORRY-Bench上的实验表明,BAIT在多个顶尖大语言模型上均保持高攻击成功率,显著优于传统越狱基线。进一步分析发现:1)以防御为导向的表述方式明显优于直接知识请求;2)细化步骤在披露升级中起关键作用;3)前两步虽有触发有害内容的可能,但极少触发过滤机制。

原文摘要 · Abstract (English)

In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.

越狱攻击模型安全推理引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。