提出首个支持多样偏好对齐的离线评估框架,提升大模型响应多样性。
Pluralistic Off-policy Evaluation and Alignment
- 构建融合效用与多样性双重目标的统一奖励函数
- 通过可分解逆倾向评分法分别评估相关性与多样性,降低方差
- 适用于需要兼顾用户偏好多样性的个性化对话系统开发
针对大语言模型在多样化人类偏好下的个性化偏好对齐需求,现有偏好对齐数据集大多来自与评估模型差异显著的策略,且现有离线策略评估方法仅关注整体效用而忽略偏好多样性。因此,将离线策略评估扩展至多元偏好对齐仍是一个开放问题。为此,我们提出首个面向大语言模型的多元离线偏好评估与对齐框架——POPE。POPE包含一个统一奖励函数,结合(1)源自人类偏好信号(如点赞或相关性评分)的协作效用分量,以及(2)受熵基覆盖率启发的多样性分量,共同反映多元对齐目标。为从记录交互中估计该奖励,我们推导出可分解的逆倾向评分(IPS)估计算法,可分别评估相关性与多样性。理论上,我们证明了该分解式IPS估计器建立了方差下界。基于离线评估的价值函数,可直接实现离线优化以进一步增强多元对齐。实验证明,POPE能有效提升响应生成的多样性,并保持模型在下游任务中的通用能力。
原文摘要 · Abstract (English)
Personalized preference alignment for LLMs with diverse human preferences requires evaluation and alignment methods that capture pluralism. Most existing preference alignment datasets are logged under policies that differ substantially from the evaluated LLMs, and existing off-policy estimators focus solely on overall utility while ignoring preference pluralism. Extending Off-Policy Evaluation (OPE) to pluralistic preference alignment, therefore, remains an open question. Thus, we propose the Pluralistic Off-Policy Evaluation (POPE), the first framework for offline pluralistic preference evaluation and alignment in LLMs. POPE includes a unified reward function that combines (1) a collaborative utility component derived from human preference signals (e.g., upvotes or relevance scores) and (2) a diversity component inspired by entropy-based coverage measures, together reflecting pluralistic alignment. Furthermore, to estimate this reward from logged interactions, we derive decomposable inverse propensity scoring (IPS) estimators that separately evaluate relevance and diversity. Theoretically, we prove that our decomposed IPS estimators establish a lower bound on their variance. With the off-policy evaluated value function, we can directly enable off-policy optimization to further enhance pluralistic alignment. Empirical results demonstrate that POPE efficiently enhances pluralistic response generation and maintains the models' general capabilities on downstream tasks
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。