通过引导模型自我推理边界,逐步诱使其泄露敏感内容。
BAIT: Boundary-Guided Disclosure Escalation via Self-Conditioned Reasoning

- 分三步诱导模型自我揭示安全边界,利用其一致性倾向
- 在多个评测集上对顶级大模型攻击成功率显著超越基线
- 适合研究模型安全漏洞或对抗攻防的学者参考
本文提出BAIT(Boundary-Aware Iterative Trap)框架,一种三步式越狱攻击方法,通过引导模型内部披露实现恶意目标。首先要求模型识别防护边界,再促使其细化该边界,最后请求具体示例。每一步均基于前一响应迭代扩展,将模型自身的推理与一致性倾向转化为披露路径。在AdvBench、JailbreakBench、AIR-Bench和SORRY-Bench上的实验表明,BAIT在多个顶尖大语言模型上均保持高攻击成功率,显著优于传统越狱基线。进一步分析发现:1)以防御为导向的表述方式明显优于直接知识请求;2)细化步骤在披露升级中起关键作用;3)前两步虽有触发有害内容的可能,但极少触发过滤机制。
原文摘要 · Abstract (English)
In this work, we propose BAIT (Boundary-Aware Iterative Trap), a three-step jailbreak framework that approaches malicious goals through internal disclosure. BAIT first asks the model to identify the protection boundary, then requires it to refine that boundary, and finally requests a detailed example. By expanding each step upon the model's previous responses, BAIT turns the model's own reasoning and consistency tendency into a disclosure pathway. Experiments on AdvBench, JailbreakBench, AIR-Bench, and SORRY-Bench demonstrate that BAIT consistently achieves strong attack success rates across top-tier large language models, significantly advancing conventional jailbreak baselines. Further analysis reveals that: 1) prevention-oriented framing significantly outperforms direct knowledge request; 2) the refinement step plays a critical role in disclosure escalation; and 3) the first two steps have a certain chance of eliciting harmful content while triggering little filtering.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。