arXiv:2607.17760cs.LGcs.AI2026-07

用多任务数据提升少样本逆强化学习的泛化与引导能力

Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning

论文配图:Generalize and Guide: Decomposing Rewards for Few-Shot Inverse Reinforcement Learning
图 1 · 摘自论文原文
  • 分离奖励为可迁移判别器和距离引导函数
  • 在复杂变化场景下平均成功率81.2%,领先基线24.7个百分点
  • 适合少样本、高变异环境下的机器人策略学习

逆强化学习(IRL)为从示范中学习提供了强大框架。然而,现实任务常存在显著自然变化(如抓取不同形状的杯子),难以收集覆盖所有情况的完整示范。实践中,虽目标任务示范有限,但往往可获得大量相关行为的异构数据集。这催生了少样本多任务示范逆强化学习(FM-IRL)问题:仅凭少量目标任务示范,结合充分的相关任务示范与在线智能体经验,学习新任务。为此,需同时恢复新任务的专家分布,并在智能体偏离时提供指导。我们提出多任务判别器近似引导逆强化学习(MPG),学习两种互补的奖励成分:(1) 可泛化的判别器,跨相关任务迁移共享结构以识别新任务中的专家行为;(2) 近似函数,衡量状态与专家行为的距离,在探索中提供纠正引导。我们在多个具有显著变化的导航与操作任务上验证方法有效性(如物体配置、桌面布局、初始机械臂姿态),平均成功率达81.2%,相比最强单任务基线平均提升24.7个百分点。

原文摘要 · Abstract (English)

Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.g., picking up mugs with varying shapes), making it impractical to collect demonstrations that fully specify a new task under every possible scenario. In practice, while demonstrations for the target task are limited, it is often easier to obtain datasets of heterogeneous but related behaviors. This motivates the problem of few-shot IRL with multi-task demonstrations (FM-IRL), where an agent must learn a new task with substantial variations from only a limited number of target-task demonstrations, together with sufficient demonstrations of related tasks and online agent experience. To do so, we must both recover the expert distribution of the new task and provide guidance when the agent deviates from it. We introduce Multitask discriminator Proximity-Guided IRL (MPG), which learns two complementary reward components: (1) a generalizable discriminator that transfers shared structure across related tasks to identify expert behavior in a new task, and (2) a proximity function that measures how far a state deviates from expert behavior and provides corrective guidance during exploration. We demonstrate the effectiveness of our method on multiple challenging navigation and manipulation tasks under significant variations (e.g., object configurations, table layouts, and initial robot poses), achieving an average success rate of 81.2%, outperforming the strongest per-task baseline by an average of 24.7 percentage points.

逆强化学习少样本学习多任务学习机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。