arXiv:2607.02755cs.LGcs.AI2026-07

让大模型在低风险任务中学会谨慎,发现它能推广到极高风险场景。

Out-of-Distribution Generalization of Risk Aversion in Language Models

  • 在低风险任务中训练模型谨慎决策,测试其在极端高风险下的表现
  • 模型在极端高风险下选择安全策略的比例达39%~70%,跨越98个数量级
  • 适用于关注AI对齐与安全机制的研究者

训练大模型在资源使用上具备风险规避倾向,可在其偏离对齐目标时提供安全防护。因只能在低风险博弈中训练风险规避,需验证其能否泛化至天文级高风险情境。为此,我们提出RiskAverseOOD基准,评估风险规避的分布外泛化能力。通过多种方法使Qwen3-8B在低风险下选择规避,发现其在极高风险下仍表现出显著风险规避:从基线2%的安全选择率,提升至SFT和绑定训练的70%、DPO的52%、激活调节的39%。另一实验中,微调后的奖励模型对风险规避推理的判别准确率达99.6%。该现象在不同规模(Qwen3-1.7B、Qwen3-14B)及模型家族(Gemma-3-12B-IT、Llama-3.1-8B-Instruct)中复现。结果表明,低风险学习的风险规避可部分泛化至极端高风险,但尚未足够稳定以作为可靠安全机制,实现一致泛化仍是开放问题。

原文摘要 · Abstract (English)

Training AIs to be risk-averse in resources could offer a failsafe in the event that AIs turn out misaligned. Misaligned but risk-averse AIs would tend to prefer low-risk, low-reward strategies like cooperation over high-risk, high-reward strategies like rebellion, limiting the downsides of any misalignment. But we can only feasibly train AIs to be risk-averse on low-stakes gambles, and we will only be safe if their risk aversion generalizes to astronomically-high-stakes gambles. Will it? To shed light on this question, we introduce RiskAverseOOD: a benchmark for measuring how well risk aversion generalizes out of distribution. We then offer some initial results. Using a variety of methods to make Qwen3-8B choose risk-aversely when the stakes are low, we find that we can induce substantial risk aversion when the stakes are astronomically high. Our models' learned risk aversion generalizes at least partially across 98 orders of magnitude. From a baseline 2% rate of choosing a safe `Cooperate' option, we see rates around 70% (SFT and tie training), 52% (DPO), and 39% (activation steering). In another experiment, our fine-tuned reward model reliably scores risk-averse reasoning above risk-neutral or excessively risk-averse alternatives (99.6% pairwise accuracy). We replicate these effects at different scales (Qwen3-1.7B and Qwen3-14B) and across model families (Gemma-3-12B-IT and Llama-3.1-8B-Instruct). Overall, we find that risk aversion learned at low stakes can generalize OOD to astronomically high stakes, though not yet consistently enough to serve as a reliable failsafe. Achieving that level of consistency is an open problem.

风险规避对齐安全泛化能力大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。