arXiv:2607.08647cs.LGcs.AI2026-07中稿 · RLC 2026

通过多环境多模态教学,让智能体学会在不同场景下都有效的奖励函数。

Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning

论文配图:Multi-Modal, Multi-Environment Machine Teaching for Robust Reward Learning
图 1 · 摘自论文原文
  • 设计分层教学算法,优先选能暴露互补约束的环境,再高效采集反馈。
  • 在相同反馈预算下,误差更低,对未见环境泛化能力更强。
  • 适合需要跨场景稳定表现的强化学习系统研发者。

随着自主智能体在多样环境中的部署,其行为需与人类意图保持一致,这就要求奖励函数具备对环境变化的鲁棒性,而非仅适应单一环境。逆向强化学习(IRL)可通过人类反馈来推断目标,但现有研究大多局限于单环境、仅示范的设定,未探讨异构反馈模态与环境动态如何共同限制可泛化的奖励函数。由于单一马尔可夫决策过程(MDP)中的示范会将奖励信息与环境结构耦合,导致学到的奖励在新环境中常失效。本文首先分析不同反馈模态对奖励的约束能力,表明在数据无限条件下,比较类反馈施加的全局约束强于其他模态。在此理论基础上,提出一种跨多个MDP运行的分层机器教学算法:先贪婪选择能揭示互补约束的环境,再在其中策略性地请求低成本反馈。实验显示,该方法在相同反馈预算下,显著降低后悔值,并在未见环境中展现出更强泛化能力,证明了多环境、多模态教学对学习动态鲁棒奖励函数的重要性。

原文摘要 · Abstract (English)

As autonomous agents are increasingly deployed across diverse operational contexts, aligning their behavior with human intent demands reward functions that remain robust to such changes rather than overfitting to any single environment. Inverse reinforcement learning (IRL) provides a principled way to infer such objectives from human feedback. However, existing analyses of optimal teaching approaches for IRL focus on single-environment, demonstration-only settings, leaving underexplored how heterogeneous feedback modalities and environment dynamics jointly constrain reward functions that generalize across multiple environments. Because demonstrations in one MDP entangle reward information with that environments specific structure, the resulting rewards frequently fail to generalize when the agent is deployed in a new setting. We first analyze how different feedback modalities constrain rewards, showing that, in the unlimited-data regime, comparisons impose strictly stronger global constraints than other modalities. Beyond this theoretical analysis, we introduce a hierarchical machine teaching algorithm for reward learning that operates across multiple MDPs. The algorithm first greedily selects informative environments that expose complementary reward constraints, then strategically queries low-cost feedback within those environments. Empirically, our method achieves substantially lower regret and stronger generalization to held-out environments than uniform teaching baselines under identical feedback budgets, demonstrating the importance of multi-environment, multi-modal teaching for learning dynamics-robust reward functions.

机器教学奖励学习多环境IRL

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。