arXiv:2501.09327cs.LGcs.AI2025-01被引 3

无需奖励信号,学习可通用的轨迹嵌入表示。

On Learning Informative Trajectory Embeddings for Imitation, Classification and Regression

  • 从状态-动作轨迹中提取多能力嵌入,不依赖奖励信号。
  • 在模仿、分类、聚类和回归任务中表现优于传统方法。
  • 嵌入具可控制行为与加性结构,适合跨领域应用。

在自动驾驶、机器人和医疗等现实序列决策任务中,从观测的状态-动作轨迹中学习对模仿、分类和聚类至关重要。例如,自动驾驶汽车需复现人类驾驶行为,而机器人与医疗系统可通过建模决策序列受益,无论数据是否来自专家。现有轨迹编码方法常针对特定任务或依赖奖励信号,限制了跨领域泛化能力。受CLIP、BERT等静态嵌入模型成功启发,我们提出一种新方法,将状态-动作轨迹嵌入到潜在空间,捕捉动态决策过程中的技能与能力。该方法无需奖励标签,支持跨领域任务的更好泛化。贡献包括:(1) 提出一种能从状态-动作数据中捕获多种能力的轨迹嵌入方法;(2) 学习到的嵌入在下游任务(模仿、分类、聚类、回归)中表现出强大表征能力;(3) 嵌入具有独特性质,如在IQ-Learn中可控制代理行为,且潜空间具加性结构。实验表明,本方法优于传统方法,为多种应用提供更灵活、强大的轨迹表示。代码已公开于https://github.com/Erasmo1015/vte。

原文摘要 · Abstract (English)

In real-world sequential decision making tasks like autonomous driving, robotics, and healthcare, learning from observed state-action trajectories is critical for tasks like imitation, classification, and clustering. For example, self-driving cars must replicate human driving behaviors, while robots and healthcare systems benefit from modeling decision sequences, whether or not they come from expert data. Existing trajectory encoding methods often focus on specific tasks or rely on reward signals, limiting their ability to generalize across domains and tasks. Inspired by the success of embedding models like CLIP and BERT in static domains, we propose a novel method for embedding state-action trajectories into a latent space that captures the skills and competencies in the dynamic underlying decision-making processes. This method operates without the need for reward labels, enabling better generalization across diverse domains and tasks. Our contributions are threefold: (1) We introduce a trajectory embedding approach that captures multiple abilities from state-action data. (2) The learned embeddings exhibit strong representational power across downstream tasks, including imitation, classification, clustering, and regression. (3) The embeddings demonstrate unique properties, such as controlling agent behaviors in IQ-Learn and an additive structure in the latent space. Experimental results confirm that our method outperforms traditional approaches, offering more flexible and powerful trajectory representations for various applications. Our code is available at https://github.com/Erasmo1015/vte.

轨迹嵌入模仿学习无监督表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。