提出因果去偏的潜在动作模型,让机器人更精准地控制动作。
Causally Debiased Latent Action Model for Embodied Action Conditioned World Models

- 通过因果去偏设计,分离动作相关与无关的视觉干扰因素。
- 在20亿和140亿参数模型上,仅用6000步微调即实现高可控性。
- 适合需要低成本动作控制的机器人学习与强化学习研究者。
动作条件世界模型(ACWMs)旨在根据具身动作模拟未来观测,为机器人规划、策略评估和数据增强提供基础。然而,学习可控制的ACWMs需要大规模带动作标签的数据,而真实世界中收集此类数据成本高昂。潜在动作模型(LAMs)通过从无标签视频中推断潜在动作缓解此瓶颈,但现有LAMs通常仅使用重建目标训练,导致动作相关动态与背景、未接触物体等无关视觉因素纠缠。本文识别出这种无关偏差是可控ACWMs的关键障碍,并引入评估指标衡量潜在动作偏差、动作跟随能力与鲁棒性。提出CD-LAM,一种基于LAM的因果去偏框架。CD-LAM引入三项高效微调目标:具身中心重建、动作中心对比学习与潜在空间校准,共同促使潜在动作表示聚焦于具身性、感知动作并避免坍缩。在20亿和140亿参数的ACWM骨干上实验表明,CD-LAM显著提升潜在动作可控性、下游机器人动作跟随、视觉保真度及适应效率,仅需6000次微调步骤,且比基线减少超过12倍的机器人动作适应更新次数。
原文摘要 · Abstract (English)
Action-conditioned world models (ACWMs) aim to simulate future observations conditioned on embodied actions, offering a promising foundation for robot planning, policy evaluation, and data augmentation. However, learning controllable ACWMs requires large-scale action-labeled data, which remains costly to collect in the real world. Latent action models (LAMs) mitigate this bottleneck by inferring latent actions from unlabeled videos, but existing LAMs are typically trained with reconstruction-only objectives and therefore entangle action-relevant dynamics with action-irrelevant visual factors such as backgrounds and untouched objects. In this work, we identify this action-irrelevant bias as a key obstacle to controllable ACWMs and introduce evaluation metrics to measure latent-action bias, action following, and robustness. We propose CD-LAM, a causally debiased framework for LAM-based ACWMs. CD-LAM introduces three efficient fine-tuning objectives: embodiment-centric reconstruction, action-centric contrastive learning, and latent space calibration, which together encourage embodiment-focused, action-aware, and calibrated non-collapsed latent action representations. Experiments on 2B and 14B ACWM backbones show that CD-LAM substantially improves latent-action controllability, downstream robot-action following, visual fidelity, and adaptation efficiency, requiring only 6k fine-tuning steps and more than 12$\times$ fewer robot-action adaptation updates than the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。