arXiv:2503.05021cs.CLcs.CR2025-03ACL被引 43

让大模型学会解释安全判断,不只是拒绝有害请求。

Safety is Not Only About Refusal: Reasoning-Enhanced Fine-tuning for Interpretable LLM Safety

  • 用推理增强微调,让模型先思考再回答
  • 在复杂场景下既能拒答又能给出合理解释
  • 适合需要可解释安全性的应用如医疗、金融

大型语言模型易受越狱攻击,传统安全对齐依赖僵硬的拒绝策略或表征工程,难以应对需语境感知的复杂安全挑战。为此,我们提出推理增强型微调框架Rational,训练模型在生成回复前进行显式安全推理。微调后的模型利用预训练阶段积累的知识,通过自生成推理实现安全内化,具备上下文敏感的决策能力。实验表明,安全不仅限于拒绝,更需结合语境的可解释、自适应响应。推理既是大模型的核心能力,也是其安全机制的基础。Rational能在复杂场景中有效拒绝有害请求,并提供有意义的上下文回应。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are vulnerable to jailbreak attacks that exploit weaknesses in traditional safety alignment, which often relies on rigid refusal heuristics or representation engineering to block harmful outputs. While they are effective for direct adversarial attacks, they fall short of broader safety challenges requiring nuanced, context-aware decision-making. To address this, we propose Reasoning-enhanced Finetuning for interpretable LLM Safety (Rational), a novel framework that trains models to engage in explicit safe reasoning before response. Fine-tuned models leverage the extensive pretraining knowledge in self-generated reasoning to bootstrap their own safety through structured reasoning, internalizing context-sensitive decision-making. Our findings suggest that safety extends beyond refusal, requiring context awareness for more robust, interpretable, and adaptive responses. Reasoning is not only a core capability of LLMs but also a fundamental mechanism for LLM safety. Rational employs reasoning-enhanced fine-tuning, allowing it to reject harmful prompts while providing meaningful and context-aware responses in complex scenarios.

大模型安全推理增强可解释性微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。