提出自监督方法学习跨机器人通用动作表示,提升少样本迁移能力。
SCAR: Self-Supervised Continuous Action Representation Learning

- 联合逆向与前向动态模型,从视觉变化中解耦动作表示。
- 在Procgen和Robotwin上实现更优的跨任务迁移与少样本适应。
- 适合研究具身智能、世界模型与动作表征的学者阅读。
尽管动作在具身智能中占据核心地位,但从视觉变化中学习可迁移的动作表征仍是根本挑战,尤其在数据有限且需跨机器人泛化时。本文认为动作不仅是辅助条件信号,更是分离可控变化与具身特异性执行的独立表征因子。为此提出SCAR:一种基于预训练生成主干的联合逆-前向动态框架,通过逆动力学模型(IDM)从隐空间观测对中推断潜在动作,再由前向动力学模型(FDM)据此预测未来动态。为确保隐空间可迁移而非通用视觉瓶颈,引入标准高斯先验正则化潜行动作后验,并采用对抗不变性抑制具身与环境相关的干扰因素。在Procgen与Robotwin数据集上的实验表明,所学统一潜动作用于世界建模时优于特定机器人的原始动作,显著提升跨机器人少样本适应与跨任务迁移性能。结果表明,动作可作为跨具身系统可控变化的共享表征,为更可迁移、泛化的世界模型提供接口。
原文摘要 · Abstract (English)
Despite the central role of action in embodied intelligence, learning transferable action representations from visual transitions remains a fundamental challenge, particularly when world models must generalize across embodiments under limited data. We argue that action is not merely an auxiliary conditioning signal, but a distinct representational factor that decouples the controllable change from embodiment-specific actuation. In this work, we propose SCAR, a joint inverse-forward dynamics framework for learning unified action representations across embodiments from visual transitions. Built on a pretrained generative backbone, SCAR uses an inverse dynamics model (IDM) to infer latent actions from latent observation pairs and a forward dynamics model (FDM) to predict future dynamics conditioned on them. To make the latent space transferable rather than a generic visual bottleneck, we regularize the latent action posterior toward a standard Gaussian prior to limit arbitrary visual encoding, and introduce adversarial invariance to suppress embodiment- and environment-specific nuisance factors. Experiments on the Procgen and Robotwin dataset show that the learned unified latent action representation serves as a stronger conditioning interface for world modeling than embodiment-specific raw actions, yielding improved cross-embodiment low-data adaptation and cross-task transfer. Taken together, these results suggest that action can be learned as a shared representation of controllable change across embodiments, providing an interface for more transferable and generalizable world models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。