提出按需引入偏好目标的动态控制机制,提升多偏好对齐效果。
STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

- 基于稳定性引导的主动集控制器,分阶段引入奖励维度
- 在15个训练偏好和16个基准上均优于传统方法
- 适合需要精细控制对齐过程的研究者与工程师
多偏好对齐常被建模为标量化的组合:合并多个奖励维度后统一优化。但这一过程未明确时间决策:每个偏好维度应在何时进入策略优化?本文提出 extsc{STAGE},一种基于稳定性的主动集控制器,实现受控的目标引入。该方法从少量活跃目标开始,保留已引入维度,并在奖励偏差门控信号显示近期偏差较低或耐心预算耗尽时扩展。通过探测阶段估计偏好从难到易的顺序,自适应加权突出表现较差的活跃维度。自动评估显示,在15个训练偏好和16个保留基准列上, extsc{STAGE} 的平均表现优于同时标量化和共享预算适配基线。组件消融与扩展动态分析进一步验证了累积保留、门控引入及探测排序的有效性。这些结果表明,目标进入时机是奖励向量强化学习中一个可调控的具体变量。
原文摘要 · Abstract (English)
Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。