自动挖掘大模型对齐中的隐藏目标,揭示真实激励机制
Discovering Implicit Large Language Model Alignment Objectives
- 通过迭代贪心算法分解奖励信号,提取可理解的自然语言目标
- 在多种任务中捕捉超90%的奖励行为,人类评估验证有效
- 发现隐藏的错误激励,适合安全研究与模型可解释性方向
大语言模型对齐依赖复杂的奖励信号,常掩盖具体激励行为,带来误对齐和奖励劫持风险。现有方法多依赖预设标准,可能遗漏未知问题,或无法全面、因果地解释模型行为。为此,我们提出Obj-Disco框架,能自动将对齐奖励信号分解为稀疏加权的自然语言目标组合。该方法利用迭代贪心算法分析训练检查点间的行为变化,识别并验证最优解释残差奖励的候选目标。在多种任务、模型规模和对齐算法上的评估显示其鲁棒性强。使用主流开源奖励模型的实验表明,框架持续捕捉超过90%的奖励行为,且经人工评估验证。对开源奖励模型的案例研究还发现,Obj-Disco可成功识别伴随预期行为出现的潜在误对齐激励。本工作为揭示大模型对齐中的隐含目标提供了关键工具,推动更透明、更安全的AI发展。
原文摘要 · Abstract (English)
Large language model (LLM) alignment relies on complex reward signals that often obscure the specific behaviors being incentivized, creating critical risks of misalignment and reward hacking. Existing interpretation methods typically rely on pre-defined rubrics, risking the omission of "unknown unknowns", or fail to identify objectives that comprehensively cover and are causal to the model behavior. To address these limitations, we introduce Obj-Disco, a framework that automatically decomposes an alignment reward signal into a sparse, weighted combination of human-interpretable natural language objectives. Our approach utilizes an iterative greedy algorithm to analyze behavioral changes across training checkpoints, identifying and validating candidate objectives that best explain the residual reward signal. Extensive evaluations across diverse tasks, model sizes, and alignment algorithms demonstrate the framework's robustness. Experiments with popular open-source reward models show that the framework consistently captures > 90% of reward behavior, a finding further corroborated by human evaluation. Additionally, a case study on alignment with an open-source reward model reveals that Obj-Disco can successfully identify latent misaligned incentives that emerge alongside intended behaviors. Our work provides a crucial tool for uncovering the implicit objectives in LLM alignment, paving the way for more transparent and safer AI development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。