无需预知行为模式数,自动识别专家不同意图并解析行为关系。
CoMI-IRL: Contrastive Multi-Intention Inverse Reinforcement Learning
- 用对比学习解耦行为表征与聚类,不依赖先验的模式数量
- 在无标签数据上表现超越现有方法,可识别未见行为
- 支持可视化分析行为间关系,适合复杂意图场景研究
逆强化学习(IRL)旨在从专家示范中推断奖励函数。当示范来自具有不同意图的多位专家时,该问题称为多意图逆强化学习(MI-IRL)。近期基于深度生成模型的MI-IRL方法将行为聚类与奖励学习耦合,但通常需要预先知道真实行为模式数 $K^*$。这种对先验知识的依赖限制了其对新行为的适应性,且仅能分析学习到的奖励,无法跨行为模式进行比较。本文提出对比多意图逆强化学习(CoMI-IRL),一种基于Transformer的无监督框架,将行为表征与聚类从下游奖励学习中解耦。实验表明,CoMI-IRL在无需 $K^*$ 先验或标签的情况下优于现有方法,同时支持行为关系的可视化分析,并可在不重新训练的前提下适应未见行为。
原文摘要 · Abstract (English)
Inverse Reinforcement Learning (IRL) seeks to infer reward functions from expert demonstrations. When demonstrations originate from multiple experts with different intentions, the problem is known as Multi-Intention IRL (MI-IRL). Recent deep generative MI-IRL approaches couple behavior clustering and reward learning, but typically require prior knowledge of the number of true behavioral modes $K^*$. This reliance on expert knowledge limits their adaptability to new behaviors, and only enables analysis related to the learned rewards, and not across the behavior modes used to train them. We propose Contrastive Multi-Intention IRL (CoMI-IRL), a transformer-based unsupervised framework that decouples behavior representation and clustering from downstream reward learning. Our experiments show that CoMI-IRL outperforms existing approaches without a priori knowledge of $K^*$ or labels, while allowing for visual interpretation of behavior relationships and adaptation to unseen behavior without full retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。