arXiv:2605.00384cs.RO2026-05被引 1

用专家混合模型提升偏好学习的鲁棒性,应对标注噪声和冲突。

PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning

论文配图:PrefMoE: Robust Preference Modeling with Mixture-of-Experts Reward Learning
图 1 · 摘自论文原文
  • 采用多专家架构与轨迹级软路由,动态组合不同偏好模式。
  • 在D4RL和MetaWorld任务上,偏好预测准确率提升,策略学习更稳定。
  • 适合处理带噪声、不一致的众包或合成标注数据,尤其适用于复杂决策场景。

基于偏好的强化学习通过对比反馈学习奖励结构,提供了一种可扩展的替代方案,避免了人工设计奖励的繁琐。然而,大规模偏好数据集(来自众包标注者或合成教师生成)常包含异质且部分冲突的监督信号,如标注者间分歧和标注者内部不一致。现有方法通常拟合单一奖励模型,迫使模型对矛盾信号取平均,限制了鲁棒性。为此,我们提出PrefMoE,一种用于鲁棒偏好建模的专家混合奖励学习框架。PrefMoE学习多个专用奖励专家,并使用轨迹级软路由自适应组合它们,从而在噪声和异质偏好监督下捕捉多样化的潜在偏好模式。负载均衡正则项进一步通过防止专家崩溃来稳定训练。在D4RL的运动基准和MetaWorld的操控任务上,PrefMoE显著提升了偏好预测的鲁棒性,并带来比强基线单模型更可靠的下游策略学习结果。

原文摘要 · Abstract (English)

Preference-based reinforcement learning offers a scalable alternative to manual reward engineering by learning reward structures from comparative feedback. However, large-scale preference datasets, whether collected from crowdsourced annotators or generated by synthetic teachers, often contain heterogeneous and partially conflicting supervision, including disagreement across annotators and inconsistency within annotators. Existing reward learning methods typically fit a single reward model to such data, forcing it to average incompatible signals and thereby limiting robustness. To solve this, we propose PrefMoE, a mixture-of-experts reward learning framework for robust preference modeling. PrefMoE learns multiple specialized reward experts and uses trajectory-level soft routing to combine them adaptively, enabling the model to capture diverse latent preference patterns under noisy and heterogeneous preference supervision. A load-balancing regularizer further stabilizes training by preventing expert collapse. Across locomotion benchmarks from D4RL and manipulation tasks from MetaWorld, PrefMoE improves preference prediction robustness and leads to more reliable downstream policy learning than strong single-model baselines.

偏好学习专家混合强化学习鲁棒建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。