从人类偏好数据中提取21个核心偏好类别,揭示人类判断的共性模式。
Learning a Canonical Basis of Human Preferences from Binary Ratings
- 从近5000种偏好中筛选出21个核心类别作为人类偏好的基础基
- 该基可解释超89%的人类偏好差异,且在不同主题上具泛化能力
- 适用于模型对齐评估与针对性微调,提升训练效率与可解释性
生成式AI的进展主要依赖于强化学习人类反馈(RLHF)等对齐技术。这类方法通常构建二元或排序选择的人类偏好数据集,并微调模型以匹配这些偏好。本文转而关注此类数据集中编码的人类偏好本质,旨在识别普遍存在的共同偏好。研究发现,从近5000种不同偏好中选出的21个偏好类别,可解释个体间超过89%的偏好差异。这一小集合类似于人类偏好的标准基,类比心理学或人脸识别中的既定发现。通过合成与实证评估,我们验证了该低秩、标准偏好基在全数据集及特定主题内均具有泛化能力。进一步证明其在模型评估中的价值:偏好类别能深入揭示模型对齐状态;在训练中,基于偏好子集的微调可有效实现模型对齐。
原文摘要 · Abstract (English)
Recent advances in generative AI have been driven by alignment techniques such as reinforcement learning from human feedback (RLHF). RLHF and related techniques typically involve constructing a dataset of binary or ranked choice human preferences and subsequently fine-tuning models to align with these preferences. This paper shifts the focus to understanding the preferences encoded in such datasets and identifying common human preferences. We find that a small subset of 21 preference categories (selected from a set of nearly 5,000 distinct preferences) captures >89% of preference variation across individuals. This small set of preferences is analogous to a canonical basis of human preferences, similar to established findings that characterize human variation in psychology or facial recognition studies. Through both synthetic and empirical evaluations, we confirm that our low-rank, canonical set of human preferences generalizes across the entire dataset and within specific topics. We further demonstrate our preference basis' utility in model evaluation, where our preference categories offer deeper insights into model alignment, and in model training, where we show that fine-tuning on preference-defined subsets successfully aligns the model accordingly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。