arXiv:2602.03825cs.LG2026-02

提出新方法让智能体在人类干预信息弱时仍能稳健学习。

The Geometry of Learning to Avoid Interventions

  • 从几何角度建模干预避免,将策略约束在可占据测度多面体的面上。
  • 发现干预信息越强,解越唯一;信息弱则存在大量次优策略。
  • 设计RIFT方法融合先验策略,提升在稀疏或弱干预下的性能。

人类干预是自主系统部署期间常见的监督来源。现有许多方法以避免干预为目标,但其后果尚不明确。本文从几何视角分析干预学习,将干预避免视为将策略限制在可占据测度多面体的一个面上。该视角揭示:干预学习的效果取决于干预策略的信息量——高信息量干预能唯一确定解,而弱干预则留下大量可行策略,其中许多为次优。针对这一不确定性,我们定义鲁棒干预学习(RIL)为在不同干预信息水平下均表现良好的学习问题。基于几何推导出残差干预微调(RIFT),该方法结合干预信号与先验策略,从可行解中筛选最优。理论证明RIFT优于先验策略,并等价于求解带诱导奖励的受限强化学习问题。实验表明,无论干预是否稀疏或信息量弱,RIFT均能持续提升策略性能。结果强调了考虑干预信息量的重要性,为从人类反馈中稳健学习提供了原则性路径。

原文摘要 · Abstract (English)

Human interventions are a common source of supervision in autonomous systems during deployment. Many existing approaches are based on avoiding interventions, yet the consequences of this objective are not well understood. We develop a geometric perspective on intervention learning that characterizes intervention avoidance as constraining policies to a face of the occupancy measure polytope. This view reveals that the effectiveness of intervention learning depends on the informativeness of the intervention strategy: highly informative interventions uniquely determine the solution, while weak interventions leave a large set of feasible policies, many of which are suboptimal. Motivated by this under-specification, we define Robust Intervention Learning (RIL) as the problem of learning policies that perform well under varying levels of intervention informativeness. From the geometric formulation, we derive Residual Intervention Fine-Tuning (RIFT), which combines interventions with a prior policy to select among feasible solutions. We show that RIFT provides provable improvement over the prior and corresponds to solving a constrained reinforcement learning problem with an induced reward. Empirically, RIFT yields consistent policy improvement across a range of intervention settings, particularly when interventions are sparse or weakly informative. These results highlight the importance of accounting for intervention informativeness and suggest a principled path toward robust learning from human feedback.

强化学习人类反馈几何建模鲁棒学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。