统一多源反馈学习奖励函数,提升智能体决策鲁棒性
MAVRL: Learning Reward Functions from Multiple Feedback Types with Amortized Variational Inference
- 将多种反馈视为贝叶斯推断中的似然,联合建模奖励函数
- 在多个控制任务中优于单反馈方法,提升对环境扰动的鲁棒性
- 自动融合不同反馈类型,无需手动调权,可解释不确定性信号
奖励学习通常依赖单一反馈类型,或通过人工加权组合多种反馈。然而,如何联合学习来自异构反馈(如示范、比较、评分、停止)的奖励函数仍不明确。本文将多类型反馈的奖励学习建模为共享潜在奖励函数的贝叶斯推断,每种反馈类型通过显式似然贡献信息。提出一种可扩展的近似变分推断方法,学习共享奖励编码器和反馈特异性似然解码器,通过优化单一证据下界进行训练。该方法避免将反馈压缩到统一中间表示,且无需手动损失平衡。在离散与连续控制基准测试中,联合推断的奖励后验优于单类型基线,有效利用反馈间的互补信息,并生成对环境扰动更鲁棒的策略。推断出的奖励不确定性还可提供模型置信度与反馈一致性分析的可解释信号。
原文摘要 · Abstract (English)
Reward learning typically relies on a single feedback type or combines multiple feedback types using manually weighted loss terms. Currently, it remains unclear how to jointly learn reward functions from heterogeneous feedback types such as demonstrations, comparisons, ratings, and stops that provide qualitatively different signals. We address this challenge by formulating reward learning from multiple feedback types as Bayesian inference over a shared latent reward function, where each feedback type contributes information through an explicit likelihood. We introduce a scalable amortized variational inference approach that learns a shared reward encoder and feedback-specific likelihood decoders and is trained by optimizing a single evidence lower bound. Our approach avoids reducing feedback to a common intermediate representation and eliminates the need for manual loss balancing. Across discrete and continuous-control benchmarks, we show that jointly inferred reward posteriors outperform single-type baselines, exploit complementary information across feedback types, and yield policies that are more robust to environment perturbations. The inferred reward uncertainty further provides interpretable signals for analyzing model confidence and consistency across feedback types.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。