arXiv:2505.17650cs.AI2025-05ACL被引 14

探究思维链能否真正降低越狱攻击危害

Does Chain-of-Thought Reasoning Really Reduce Harmfulness from Jailbreaking?

  • 通过理论分析揭示思维链的双重作用机制
  • 提出新越狱方法FicDetail并验证理论发现
  • 适合关注大模型安全与推理机制的研究者

越狱攻击在近期采用思维链(CoT)推理的模型上表现不佳,但其内在机理尚不明确,仅依赖推理能力可能带来安全隐患。本文旨在回答:思维链推理是否真的能降低越狱攻击的危害?通过严谨的理论分析,我们证明了思维链对越狱危害具有双重影响。基于这些理论洞见,我们提出一种新型越狱方法FicDetail,其实验性能验证了理论结论。

原文摘要 · Abstract (English)

Jailbreak attacks have been observed to largely fail against recent reasoning models enhanced by Chain-of-Thought (CoT) reasoning. However, the underlying mechanism remains underexplored, and relying solely on reasoning capacity may raise security concerns. In this paper, we try to answer the question: Does CoT reasoning really reduce harmfulness from jailbreaking? Through rigorous theoretical analysis, we demonstrate that CoT reasoning has dual effects on jailbreaking harmfulness. Based on the theoretical insights, we propose a novel jailbreak method, FicDetail, whose practical performance validates our theoretical findings.

大模型安全越狱攻击思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。