MiCRo通过混合建模与动态路由,让大模型更好理解多样化的个人偏好。
MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
- 用混合模型捕捉不同人群的偏好分布,突破单一奖励函数限制。
- 在多个数据集上显著提升个性化任务表现,效果优于现有方法。
- 无需细粒度标注,适合需要高效个性化对齐的场景。
基于人类反馈的强化学习(RLHF)中,奖励建模是构建安全基础模型的关键步骤。然而,传统的布拉德利-特里(BT)模型假设存在全局奖励函数,无法捕捉人类偏好的内在多样性与异质性,导致大语言模型难以实现个性化和多元对齐。理论上,当人类偏好服从多个子群体的混合分布时,单个BT模型存在不可消除的误差。现有方法如多目标学习虽能缓解问题,但依赖细粒度标注,成本高且受限于预定义属性,难以充分表达人类价值观的丰富性。为此,本文提出MiCRo,一种两阶段框架,利用大规模二元偏好数据集,在无需显式细粒度标注的前提下,增强个性化偏好学习。第一阶段引入上下文感知的混合建模方法,识别多样化偏好;第二阶段采用在线路由策略,根据具体上下文动态调整混合权重,解决歧义问题,实现高效可扩展的偏好适应,仅需极少额外监督。多数据集实验表明,MiCRo有效捕捉多样偏好,显著提升下游个性化性能。
原文摘要 · Abstract (English)
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently diverse and heterogeneous human preferences. Hence, such oversimplification limits LLMs from supporting personalization and pluralistic alignment. Theoretically, we show that when human preferences follow a mixture distribution of diverse subgroups, a single BT model has an irreducible error. While existing solutions, such as multi-objective learning with fine-grained annotations, help address this issue, they are costly and constrained by predefined attributes, failing to fully capture the richness of human values. In this work, we introduce MiCRo, a two-stage framework that enhances personalized preference learning by leveraging large-scale binary preference datasets without requiring explicit fine-grained annotations. In the first stage, MiCRo introduces context-aware mixture modeling approach to capture diverse human preferences. In the second stage, MiCRo integrates an online routing strategy that dynamically adapts mixture weights based on specific context to resolve ambiguity, allowing for efficient and scalable preference adaptation with minimal additional supervision. Experiments on multiple preference datasets demonstrate that MiCRo effectively captures diverse human preferences and significantly improves downstream personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。