arXiv:2601.18751cs.LGcs.AI2026-01

让智能体自动识别并应对不同专家反馈的可信度,提升强化学习鲁棒性。

Trust, Don't Trust, or Flip: Robust Preference-Based Reinforcement Learning with Multi-Expert Feedback

  • 通过联合学习奖励模型与专家信任参数,动态判断反馈可信度。
  • 在对抗性标注下仍保持接近最优性能,优于现有方法。
  • 无需额外特征,适配现有强化学习流程,适合真实场景应用。

基于偏好强化学习(PBRL)通过成对轨迹比较学习,替代显式奖励设计。然而,现实中的偏好数据来自可靠性各异的标注者:部分准确,部分嘈杂,部分系统性地提供错误偏好。现有方法或均等对待所有反馈,或尝试过滤不可靠来源,但在面对系统性误导的标注者时失效。我们提出TriTrust-PBRL(TTP),一个统一框架,从多专家偏好反馈中联合学习共享奖励模型与专家特定的信任参数。核心思想是信任参数在梯度优化过程中自然演化为正(信任)、近零(忽略)或负(反转),使模型能自动纠正对抗性偏好,而非简单丢弃污染数据。理论分析证明了可辨识性,并揭示了无需显式监督即可自然形成专家分离的梯度机制。我们在四个不同领域(元世界操控任务、DM控制运动任务)评估TTP,涵盖多种数据污染场景。TTP实现当前最优鲁棒性,在对抗性污染下维持近似最优性能,而标准PBRL方法则严重崩溃。特别地,TTP成功从包含可靠与对抗性标注者的混合专家池中学习,仅需专家标识索引,无需额外特征,且可无缝集成至现有PBRL流程。

原文摘要 · Abstract (English)

Preference-based reinforcement learning (PBRL) offers a promising alternative to explicit reward engineering by learning from pairwise trajectory comparisons. However, real-world preference data often comes from heterogeneous annotators with varying reliability; some accurate, some noisy, and some systematically adversarial. Existing PBRL methods either treat all feedback equally or attempt to filter out unreliable sources, but both approaches fail when faced with adversarial annotators who systematically provide incorrect preferences. We introduce TriTrust-PBRL (TTP), a unified framework that jointly learns a shared reward model and expert-specific trust parameters from multi-expert preference feedback. The key insight is that trust parameters naturally evolve during gradient-based optimization to be positive (trust), near zero (ignore), or negative (flip), enabling the model to automatically invert adversarial preferences and recover useful signal rather than merely discarding corrupted feedback. We provide theoretical analysis establishing identifiability guarantees and detailed gradient analysis that explains how expert separation emerges naturally during training without explicit supervision. Empirically, we evaluate TTP on four diverse domains spanning manipulation tasks (MetaWorld) and locomotion (DM Control) under various corruption scenarios. TTP achieves state-of-the-art robustness, maintaining near-oracle performance under adversarial corruption while standard PBRL methods fail catastrophically. Notably, TTP outperforms existing baselines by successfully learning from mixed expert pools containing both reliable and adversarial annotators, all while requiring no expert features beyond identification indices and integrating seamlessly with existing PBRL pipelines.

强化学习偏好学习鲁棒性多专家

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。