不依赖最优示范者,通过多源次优数据学习可靠奖励函数
Inverse Reinforcement Learning without an Optimal Demonstrator: A Feasible Reward Set Approach
- 用线性约束表示每个示范者的次优程度,交集生成可行奖励集
- 新增示范者可严格收紧可行集,理论保证在无近优示范时仍能恢复真实奖励
- 适用于低质量、多样化示范数据场景,尤其适合大模型微调等实际应用
逆强化学习通常假设示范来自单一最优行为者,但现实中数据常来自多个次优程度各异的示范者。本文提出可行奖励集框架:对每位示范者,将其声明的次优水平转化为线性约束,再通过交集整合所有示范者信息。理论分析表明,随着数据增加,联合可行集单调缩小,并给出了新示范者严格收紧集合的精确条件。我们建立了两项恢复保证:其一依赖于接近最优状态分布,其二仅需充分覆盖且无需近优示范者。实践中,提出解决所得奖励集固有模糊性的策略,并设计了支持函数逼近的离线算法,适用于高维环境。在小规模网格世界和大型语言模型(LLM)微调场景的实验验证了理论预测,且性能优于基线方法。
原文摘要 · Abstract (English)
Inverse reinforcement learning (IRL) typically assumes demonstrations from a single optimal demonstrator, but in many applications data come from multiple imperfect demonstrators with heterogeneous suboptimality levels. We study reward learning in this setting through a feasible-reward-set framework: for each demonstrator, we encode its declared suboptimality level as a linear constraint and intersect the resulting feasible sets across demonstrators. Our theoretical analysis shows that the joint feasible set shrinks monotonically as data are added, and we give an exact characterization of when a new demonstrator strictly tightens it. We further establish two recovery guarantees for the feasible reward set of the ground-truth optimal demonstrator: one bound depends on closeness to the optimal occupancy, while the other requires only sufficient coverage and no near-optimal demonstrator. On the practical side, we introduce strategies to address the inherent reward ambiguity in the obtained reward set and provide an offline algorithm with function approximation for high-dimensional environments. Experiments in tabular grid-world and large language model (LLM) fine-tuning settings are consistent with the theoretical predictions and demonstrate the effectiveness of the proposed framework over baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。