arXiv:2603.15397cs.CRcs.AI2026-03

实时检测并修正大模型推理中的安全风险,显著降低越狱攻击成功率。

SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration

  • 在推理全过程动态评估安全风险,非仅检查最终输出。
  • 将越狱攻击成功率从58.97%降至12.31%,性能基本不受影响。
  • 适合需要高安全性的复杂推理场景,如医疗、金融决策支持。

大型语言模型在复杂推理任务中表现出色,但仍易受越狱攻击影响,导致安全对齐失效。现有防御机制多依赖对最终输出的后处理过滤,未监控中间推理步骤,易被恶意操纵。为此,本文提出一种更安全的思维链框架SFCoT,通过三重安全评分系统与多视角一致性验证机制,在推理过程中实时识别潜在风险,并由动态干预模块进行针对性校正,引导推理走向安全结果。实验表明,SFCoT将攻击成功率从58.97%降至12.31%,有效提升模型安全性,且未显著牺牲通用性能。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities in complex reasoning tasks. However, they remain highly susceptible to jailbreak attacks that undermine their safety alignment. Existing defense mechanisms typically rely on post hoc filtering applied only to the final output, leaving intermediate reasoning steps unmonitored and vulnerable to adversarial manipulation. To address this gap, this paper proposes a SaFer Chain-of-Thought (SFCoT) framework, which proactively evaluates and calibrates potentially unsafe reasoning steps in real time. SFCoT incorporates a three-tier safety scoring system alongside a multi-perspective consistency verification mechanism, designed to detect potential risks throughout the reasoning process. A dynamic intervention module subsequently performs targeted calibration to redirect reasoning trajectories toward safe outcomes. Experimental results demonstrate that SFCoT reduces the attack success rate from $58.97\%$ to $12.31\%$, demonstrating it as an effective and efficient LLM safety enhancement method without a significant decline in general performance.

大模型安全越狱防御推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。