用压缩状态表示让机器人学会通用运动,无需标注数据
StaMo: Unsupervised Learning of Generalizable Robot Motion from Compact State Representation
- 用轻量编码器+扩散变换器解码器,学出两个令牌的紧凑状态表征
- 在LIBERO上性能提升14.3%,真实任务成功率提高30%,推理开销极低
- 状态差值自动成为有效隐动作,适合无监督机器人学习与政策协同训练
具身智能的核心挑战在于构建高效且信息丰富的状态表征以实现世界建模与决策。现有方法常难以兼顾简洁性与任务关键信息,导致表征冗余或缺失。本文提出一种无监督方法,通过轻量编码器与预训练扩散变换器(DiT)解码器,学习高度压缩的双令牌状态表示,利用其强大的生成先验。该表示兼具高效、可解释性,并可无缝集成至现有视觉语言动作模型(VLA),在LIBERO上提升性能14.3%,真实任务成功率提升30%,且推理开销极小。更重要的是,通过潜在插值获得的两令牌差异,自然形成高效隐动作,可进一步解码为可执行机器人动作。这一涌现能力表明,该表示捕捉了结构化动态而无需显式监督。我们将其命名为StaMo,即从静态图像编码的紧凑状态中学习通用机器人运动,挑战了依赖复杂架构和视频数据学习隐动作的主流范式。所得隐动作还提升了策略协同训练效果,在基准上优于先前方法10.4%,并具备更好可解释性。该方法在真实机器人数据、仿真环境及人类第一视角视频等多种数据源上均表现良好。
原文摘要 · Abstract (English)
A fundamental challenge in embodied intelligence is developing expressive and compact state representations for efficient world modeling and decision making. However, existing methods often fail to achieve this balance, yielding representations that are either overly redundant or lacking in task-critical information. We propose an unsupervised approach that learns a highly compressed two-token state representation using a lightweight encoder and a pre-trained Diffusion Transformer (DiT) decoder, capitalizing on its strong generative prior. Our representation is efficient, interpretable, and integrates seamlessly into existing VLA-based models, improving performance by 14.3% on LIBERO and 30% in real-world task success with minimal inference overhead. More importantly, we find that the difference between these tokens, obtained via latent interpolation, naturally serves as a highly effective latent action, which can be further decoded into executable robot actions. This emergent capability reveals that our representation captures structured dynamics without explicit supervision. We name our method StaMo for its ability to learn generalizable robotic Motion from compact State representation, which is encoded from static images, challenging the prevalent dependence to learning latent action on complex architectures and video data. The resulting latent actions also enhance policy co-training, outperforming prior methods by 10.4% with improved interpretability. Moreover, our approach scales effectively across diverse data sources, including real-world robot data, simulation, and human egocentric video.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。