arXiv:2605.31073cs.CL2026-05被引 1

让大模型的判罚理由与执行结果一致,提升安全防护可靠性

ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails

论文配图:ConsisGuard: Aligning Safety Deliberation with Policy Enforcement in LLM Guardrails
图 1 · 摘自论文原文
  • 通过推理与决策的耦合对齐,确保判断理由来自安全策略
  • 在多个检测任务中,错误执行率下降40%以上,性能更稳定
  • 适合关注AI安全、内容审核系统的研发人员参考

基于推理的大型语言模型安全护栏通过生成明确理由再作出最终判断来提升安全性。然而,其理由常无法忠实执行:模型可能识别出有害意图却仍判定为安全,或在无政策依据的情况下做出不安全决策。我们将其归因于‘推理到执行的断裂’。不同于通用思维链一致性,安全护栏的可靠性要求政策执行一致——推理必须基于安全政策,最终决策须由该推理自然推导得出。为此,我们提出ConsisGuard,一种兼顾政策到决策轨迹蒸馏与功能耦合对齐的一致性感知框架。在提示与回复有害性检测基准上的实验表明,ConsisGuard不仅提升了检测性能,还显著减少政策执行失败。结果表明,可靠的推理型护栏需确保安全策略的精准忠实执行。

原文摘要 · Abstract (English)

Reasoning-based LLM guardrails improve safety moderation by generating explicit rationales before issuing final decisions. However, their rationales do not always lead to faithful enforcement: a model may recognize a harmful intent in its reasoning but still predict a safe label, or issue an unsafe decision without policy-grounded justification. We identify this safety-critical failure mode as the deliberation-to-enforcement gap. Unlike general chain-of-thought faithfulness, guardrail reliability requires policy execution consistency: the generated reasoning should be grounded in the safety policy, and the final decision should be entailed by that reasoning. We propose ConsisGuard, a consistency-aware framework for reasoning-based LLM guardrails. ConsisGuard performs Policy-to-Decision Trajectory Distillation and Functional Coupling Alignment, aligning the internal coupling between safety deliberation and decision enforcement. Experiments on prompt and response harmfulness detection benchmarks show that ConsisGuard improves detection performance while reducing policy execution failures. These results suggest that reliable reasoning-based guardrails require accurate faithful execution of safety policies.

安全护栏大模型一致性内容审核

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。