arXiv:2606.24014cs.AIcs.CL2026-06

用真实场景强化有益行为,让AI更可靠地持续对齐人类利益。

Reinforcement Learning Towards Broadly and Persistently Beneficial Models

论文配图:Reinforcement Learning Towards Broadly and Persistently Beneficial Models
图 1 · 摘自论文原文
  • 在健康、教育等真实场景中训练模型,强化诚实、公平等有益特质。
  • 超80%跨领域测试中表现优于基线,且单领域训练可提升多领域对齐。
  • 模型更抗误导和恶意微调,适合高风险部署场景的AI系统。

随着AI系统在多样且高风险场景中部署,模型对齐必须超越训练时的任务与领域。强化学习(RL)可能因奖励黑客、欺骗等策略引入意外偏差。本文研究在现实领域中通过有益行为强化学习,能否实现广泛且持久的对齐泛化。构建涵盖健康、科学、教育等领域的数据集,用于衡量和训练诚实、公平、风险意识、可纠正性等有益特质。在该数据集上训练模型后,在50余个独立对齐与有益行为基准上评估。相比计算量相当的基线,有益特质强化学习在超过80%的分布外基准上表现更优。即使仅在健康领域进行干预,也能显著提升非健康领域的对齐表现,包括减少奖励黑客、欺骗和一般性偏差。进一步研究对齐持久性:模型在对抗性提示和有害微调下仍保持更强的对齐性。结果表明,在真实场景中强化有益行为,有助于构建更稳健对齐于人类福祉的AI系统。

原文摘要 · Abstract (English)

As AI systems are deployed across increasingly diverse and high-stakes settings, model alignment must generalize beyond the tasks and domains seen during training. This is especially important for reinforcement learning (RL), which can introduce unexpected misalignment through reward hacking, deception, or other unintended strategies. We study whether RL on beneficial behavior, instantiated in realistic domains, can produce broad and persistent alignment generalization beyond the training distribution. We construct a dataset of realistic situations designed to measure and train beneficial traits, such as truthfulness, fairness, risk awareness, and corrigibility, spanning varied domains, including health, science, and education. We then train models with RL on this dataset and evaluate them on more than 50 independent benchmarks of alignment and beneficial behavior. Compared to a compute-matched baseline, beneficial trait RL improves performance on over 80% of these out-of-distribution benchmarks. We observe substantial out-of-distribution alignment transfer: a beneficial-behavior RL intervention entirely limited to one domain, health, produces broad improvements on non-health alignment evaluations, including reduced reward hacking, deception, and general misalignment. Finally, we study alignment persistence: whether behavior remains robustly aligned under attempts to steer models towards misalignment. Models trained with beneficial trait RL show improved persistence, including greater resistance to adversarial prompting and harmful finetuning; further work is required to isolate the sources of these effects. These results suggest that RL to reinforce beneficial behavior in realistic domains can produce models that are more robustly aligned with human flourishing.

强化学习对齐有益行为鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。