用稀疏专家混合模型实现个性化偏好建模,让每类偏好由专用专家处理。
Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling

- 通过稀疏路由机制让不同专家专注处理特定偏好类型。
- 在真实数据上提升个性化推荐效果,专家权重可解释适应过程。
- 无需额外标注,适合需要可解释个性化推荐的场景。
偏好建模在人类反馈强化学习(RLHF)中起核心作用,使大语言模型(LLMs)与人类价值观对齐。然而,现有方法多假设存在统一奖励函数,忽视了人类偏好的多样性和异质性。为解决此问题且不增加标注成本,近期工作尝试从二元数据中学习多个偏好成分并组合以建模个体偏好。但这些成分常无法捕捉连贯、解耦的模式,限制了其可解释性与个性化效果。本文提出一种稀疏混合专家(Sparse MoE)奖励模型,在二元偏好数据训练中鼓励稀疏路由与专家多样性。在控制实验与真实世界实验中,该模型学习到可解释的路由模式和专业化专家,提升了测试时个性化表现;后适应阶段的专家权重变化提供了分析模型如何适配个性化偏好的定性视角。
原文摘要 · Abstract (English)
Preference modeling plays a central role in reinforcement learning from human feedback (RLHF), enabling large language models (LLMs) to align with human values. However, most existing approaches assume a universal reward function, neglecting the diversity and heterogeneity of human preferences. To address this limitation without additional annotation costs, recent work has proposed learning multiple preference components from binary data and combining them to model individual preferences. Nevertheless, these components often fail to capture coherent and disentangled patterns, limiting their interpretability and effectiveness for personalization. In this work, we propose a sparse Mixture-of-Experts (MoE) reward model that encourages sparse routing and expert diversity during training on binary preference data. Across controlled and real-world experiments, sparse MoE learns interpretable routing patterns and specialized experts. It also improves test-time personalization, and post-adaptation shifts in expert weights provide a qualitative lens for analyzing how the model adapts to personalized preferences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。