用0.5亿参数模型实现高效机器人操作,自监督学习动作相关潜在表示。
SLIM-0.5B: Learning Action-Grounded Predictive Latents for Robot Manipulation

- 自监督掩码轨迹预测,联合学习动作重建与未来状态预测
- 仅0.5B参数即在仿真和真实场景中达到主流大模型水平
- 适合资源受限的实时机器人控制任务,推理快且显存占用低
视觉-语言-动作策略依赖大型多模态骨干网络,在每个控制步骤中同时完成感知、语言条件化和动作生成。其中大量计算能力用于支持开放域语义,而连续机器人操作主要需要对观测、动作及动作引发的转换进行紧凑表征。像素级世界模型是另一路径,但预测与控制无关的视觉细节会带来不必要的开销。本文提出SLIM(自监督潜在交互模型),一个0.5B参数的紧凑型潜在交互策略。SLIM通过自监督掩码轨迹预测学习动作相关的预测潜在表示,同时捕捉动作条件下的未来转移以及解释观测变化的动作。该模型采用混合变换器(MoT)架构建模观测潜在与动作标记间的交互,并使用流匹配训练语言条件动作生成。在仿真基准和真实世界评估中,SLIM以更少参数、无需额外具身预训练、更低推理延迟和显著降低的GPU内存占用,达到或超越代表性大规模VLA和世界-动作模型基线。
原文摘要 · Abstract (English)
Vision-language-action policies rely on large multimodal backbones to jointly perform perception, language conditioning, and action generation at every control step. Much of this capacity supports open-domain semantics, whereas continuous robot manipulation primarily requires compact representations of observations, actions, and the transitions induced by actions. Pixel-level world models provide another route, but predicting visual details irrelevant to control can be unnecessarily expensive. We propose SLIM (Self-supervised Latent Interaction Model), a compact 0.5B-parameter latent interaction policy. SLIM learns action-grounded predictive latents that capture both action-conditioned future transitions and the actions that explain observed changes. SLIM learns these representations through self-supervised masked trajectory prediction, combining action reconstruction with future-latent prediction. A compact Mixture-of-Transformers (MoT) backbone models interactions between observation latents and action tokens. The resulting policy is trained with flow matching for language-conditioned action generation. Across simulation benchmarks and real-world evaluation, SLIM matches or exceeds representative large-scale VLA and world-action-model baselines with fewer parameters, no additional embodied pretraining, lower inference latency, and substantially lower GPU memory usage.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。