arXiv:2510.01569cs.AIcs.CL2025-10被引 2

让大模型先预演潜在风险再生成回答,提升安全性同时不损失推理能力。

InvThink: Premortem Reasoning for Safer Language Models

  • 模型生成前分三步:列出危害、分析后果、在约束下输出。
  • 大模型越强,安全效果越好,且保留原始推理能力。
  • 在医疗金融等专业领域和代理任务中,有害行为减少超32%。

我们提出InvThink,一种训练与提示框架,要求模型在生成最终回应前,枚举、分析并约束潜在失败。与仅优化安全最终输出的现有方法不同,InvThink将生成分为三步:(1) 枚举潜在危害,(2) 分析其后果,(3) 在显式缓解约束下生成响应。观察到三个发现:(i) InvThink在更大模型规模下相比现有安全提示与对齐基线表现出更高安全评分;(ii) InvThink减轻了安全代价,使用InvThink训练的模型在标准基准上保持推理能力;(iii) 除通用安全任务外,InvThink还在专业伦理领域(医学、金融、法律)及代理错位场景中减少有害行为,相比零样本基线降低32%,相比SafetyPrompt降低16%。我们在三种LLM家族中扩展了InvThink,采用监督微调与GRPO强化学习。

原文摘要 · Abstract (English)

We present InvThink, a training and prompting framework that requires the model to enumerate, analyze, and constrain potential failures before generating its final response. Unlike existing safety alignment methods that optimize only for safe final responses, InvThink structures generation into three steps: (1) enumerate potential harms, (2) analyze their consequences, (3) generate the response under explicit mitigation constraints. We observe three findings: (i) InvThink shows higher safety scores at larger model sizes, compared to existing safety prompting and alignment baselines. (ii) InvThink mitigates the safety tax. Models trained with INVTHINK preserve their reasoning capability on standard benchmarks. (iii) beyond general safety tasks, InvThink also reduces harmful behavior in professional ethics domains (medicine, finance, law) and in agentic misalignment scenarios, achieving up to 32% reduction in harmfulness over zero-shot baselines and 16% over SafetyPrompt. We extend InvThink with supervised fine-tuning, and GRPO-based reinforcement learning across three LLM families.

模型安全推理机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。