用柯尔莫哥洛夫算子实现机器人视觉表征的高效线性化,提升控制稳定性。
RoboKoop: Efficient Control Conditioned Representations from Visual Input in Robotics using Koopman Operator
- 基于柯尔莫哥洛夫理论,从视觉输入中学习任务相关的线性化表征。
- 在高维潜在空间中提取表征,使线性控制器在长时序上更稳定准确。
- 适合需要高效视觉-控制融合的机器人研究者,尤其关注控制稳定性。
构建能从高维观测中执行复杂控制任务的智能体,需具备稳健的任务控制策略,并将底层视觉表征适配至具体任务。现有方法多依赖大量训练样本,采用两阶段学习:先预训练视觉模型,再在其上学习控制器。本文从柯尔莫哥洛夫(Koopman)理论出发,学习特定下游任务条件下的视觉表征,用于代理的稳定控制学习。提出对比谱柯尔莫哥洛夫嵌入网络(Contrastive Spectral Koopman Embedding),可从机器人视觉数据中高效提取高维潜在空间中的线性化表征,并利用强化学习在这些表征上实现离策略控制,采用线性控制器。该方法显著提升了梯度动力学下的长期稳定性,相较于现有方法,在长时域任务中大幅提高了策略学习的效率与准确性。
原文摘要 · Abstract (English)
Developing agents that can perform complex control tasks from high-dimensional observations is a core ability of autonomous agents that requires underlying robust task control policies and adapting the underlying visual representations to the task. Most existing policies need a lot of training samples and treat this problem from the lens of two-stage learning with a controller learned on top of pre-trained vision models. We approach this problem from the lens of Koopman theory and learn visual representations from robotic agents conditioned on specific downstream tasks in the context of learning stabilizing control for the agent. We introduce a Contrastive Spectral Koopman Embedding network that allows us to learn efficient linearized visual representations from the agent's visual data in a high dimensional latent space and utilizes reinforcement learning to perform off-policy control on top of the extracted representations with a linear controller. Our method enhances stability and control in gradient dynamics over time, significantly outperforming existing approaches by improving efficiency and accuracy in learning task policies over extended horizons.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。