arXiv:2512.10016cs.LG2025-12被引 3

用无标签轨迹训练控制模型,仅需少量带动作标签数据即可高效学习。

Latent Action World Models for Control with Unlabeled Trajectories

  • 构建共享潜空间,统一处理带动作和无动作的轨迹数据。
  • 在DeepMind Control Suite上,动作标签数据减少约90%仍达强性能。
  • 适合缺乏标注数据的离线强化学习场景,可融合被动观察与主动交互。

受人类结合直接操作与被动观察(如看视频)的启发,我们研究从异构数据中学习的世界模型。传统世界模型依赖带动作标签的轨迹,当动作标签稀缺时效果受限。本文提出一类潜动作世界模型,通过学习共享的潜动作表示,联合使用带动作和无动作的数据。该潜空间将观测到的控制信号与从被动观察中推断的动作对齐,使单一动态模型能在大规模无标签轨迹上训练,仅需少量带标签样本。利用该模型通过离线强化学习学习潜动作策略,打通了传统上分离的离线强化学习(依赖带动作数据)与无动作训练(极少用于后续强化学习)领域。在DeepMind Control Suite上,本方法性能强劲,同时使用的动作标签样本比纯带动作基线少约一个数量级。结果表明,潜动作使世界模型能有效利用被动与主动数据,显著提升学习效率。

原文摘要 · Abstract (English)

Inspired by how humans combine direct interaction with action-free experience (e.g., videos), we study world models that learn from heterogeneous data. Standard world models typically rely on action-conditioned trajectories, which limits effectiveness when action labels are scarce. We introduce a family of latent-action world models that jointly use action-conditioned and action-free data by learning a shared latent action representation. This latent space aligns observed control signals with actions inferred from passive observations, enabling a single dynamics model to train on large-scale unlabeled trajectories while requiring only a small set of action-labeled ones. We use the latent-action world model to learn a latent-action policy through offline reinforcement learning (RL), thereby bridging two traditionally separate domains: offline RL, which typically relies on action-conditioned data, and action-free training, which is rarely used with subsequent RL. On the DeepMind Control Suite, our approach achieves strong performance while using about an order of magnitude fewer action-labeled samples than purely action-conditioned baselines. These results show that latent actions enable training on both passive and interactive data, which makes world models learn more efficiently.

世界模型离线RL潜空间无标签数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。