通过逆博弈框架,让机器人学会与人协作的复杂互动行为。
Structured Imitation Learning of Interactive Policies through Inverse Games
- 用生成模型学习人类示范中的个体行为模式。
- 通过求解逆博弈问题,捕捉多智能体间的依赖关系。
- 仅用50次示范就接近真实交互策略表现,适合人机协作场景。
基于生成模型的模仿学习方法在从人类示范中学习高复杂度运动技能方面已取得显著成果。然而,在无明确沟通的情况下,让智能体在共享空间中与人类协调互动仍具挑战性,因为多智能体交互的行為复杂度远高于非交互任务。本文提出一种结构化模仿学习框架,将生成式单智能体策略学习与灵活且表达力强的博弈论结构相结合。方法分两步:首先,利用标准模仿学习从多智能体示范中学习个体行为模式;其次,通过求解逆博弈问题,结构化地学习智能体间的依赖关系。在一个人造的五智能体社交导航任务中,初步结果显示,该方法显著优于非交互策略,并仅用50次示范即可达到与真实交互策略相当的性能。这凸显了结构化模仿学习在交互场景中的潜力。
原文摘要 · Abstract (English)
Generative model-based imitation learning methods have recently achieved strong results in learning high-complexity motor skills from human demonstrations. However, imitation learning of interactive policies that coordinate with humans in shared spaces without explicit communication remains challenging, due to the significantly higher behavioral complexity in multi-agent interactions compared to non-interactive tasks. In this work, we introduce a structured imitation learning framework for interactive policies by combining generative single-agent policy learning with a flexible yet expressive game-theoretic structure. Our method explicitly separates learning into two steps: first, we learn individual behavioral patterns from multi-agent demonstrations using standard imitation learning; then, we structurally learn inter-agent dependencies by solving an inverse game problem. Preliminary results in a synthetic 5-agent social navigation task show that our method significantly improves non-interactive policies and performs comparably to the ground truth interactive policy using only 50 demonstrations. These results highlight the potential of structured imitation learning in interactive settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。