用连续动作表示跨环境建模,无需标签快速适应新场景
Inter-environmental world modeling for continuous and compositional dynamics
- 基于李群理论与对象中心自编码器,学习连续动作表征
- 仅用视频帧即可训练,在合成与真实数据上零/少标签快速适配
- 适合需要跨环境泛化与低标签依赖的机器人控制研究
当前多数世界模型基于自回归框架,依赖离散的动作与观测表示,虽在目标环境建模中表现良好,但难以跨环境泛化。人类能整合多环境经验进行心理模拟与控制学习,受此启发,本文提出无监督框架WLA(World modeling through Lie Action),通过李群理论与对象中心自编码器,联合建模多个环境的动力学,学习连续潜空间动作表示。该框架构建了高可控性与强预测能力的控制接口。在合成基准与真实世界数据集上,WLA仅需视频帧输入,无需或极少动作标签,即可快速适应包含新动作集的新环境。
原文摘要 · Abstract (English)
Various world model frameworks are being developed today based on autoregressive frameworks that rely on discrete representations of actions and observations, and these frameworks are succeeding in constructing interactive generative models for the target environment of interest. Meanwhile, humans demonstrate remarkable generalization abilities to combine experiences in multiple environments to mentally simulate and learn to control agents in diverse environments. Inspired by this human capability, we introduce World modeling through Lie Action (WLA), an unsupervised framework that learns continuous latent action representations to simulate across environments. WLA learns a control interface with high controllability and predictive ability by simultaneously modeling the dynamics of multiple environments using Lie group theory and object-centric autoencoder. On synthetic benchmark and real-world datasets, we demonstrate that WLA can be trained using only video frames and, with minimal or no action labels, can quickly adapt to new environments with novel action sets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。