arXiv:2607.04652cs.RO2026-07被引 1

用冻结的视频世界模型生成机器人操作的方向性提示,提升少样本学习效果。

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation

论文配图:KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation
图 1 · 摘自论文原文
  • 从冻结的潜空间视频模型中提取单步速度图作为交互方向线索
  • 在LIBERO上达90.6%成功率,在RoboTwin2.0易/难场景分别达65.7%/22.4%
  • 无需微调或滚动预测,适合少样本机器人操控任务

从少量示范中学习操作需要捕捉交互位置与起始方式的视觉先验;静态先验如分割掩码仅包含位置信息。我们提出KAM-WM框架,通过查询一次流匹配图像到视频主干,将单步潜空间速度解释为运动学可操作图(KAM),提供任务相关的交互区域与粗略运动结构。轻量级Perceiver将KAM压缩为条件令牌,与RGB观测和本体感知共同指导扩散策略。在LIBERO和RoboTwin2.0上,KAM-WM平均成功率达90.6%,在RoboTwin2.0的Easy与Hard设置下分别取得65.7%和22.4%的成功率。对比零阶掩码先验的受控实验表明,部分性能提升源于超越空间定位的方向性信息。结果表明,在评估设置下,冻结的视频模型可无需测试时未来滚动开销地提供有效的一阶视觉先验。

原文摘要 · Abstract (English)

Learning manipulation from few demonstrations requires visual priors that capture not only where to interact, but also how the interaction should begin; static priors such as segmentation masks encode only the former. We present KAM-WM, a framework that extracts a coarse directional interaction cue from a frozen latent video world model without rollout or world-model fine-tuning. KAM-WM queries a Flow Matching image-to-video backbone once and interprets its single-step latent velocity as a Kinematic Affordance Map (KAM), which provides task-conditioned interaction regions and coarse motion structure. A lightweight Perceiver compresses KAM into tokens that condition a diffusion policy together with RGB observations and proprioception. Across LIBERO and RoboTwin2.0, KAM-WM reaches 90.6% average success on LIBERO and achieves 65.7% and 22.4% success rates in the Easy and Hard settings on RoboTwin2.0, respectively. Controlled comparisons against a zero-order mask prior suggest that part of the gains comes from directional information beyond spatial localization alone. These results indicate that, in the evaluated settings, a frozen video model can provide a useful first-order visual prior for control without the test-time cost of future rollout.

机器人操控视觉先验扩散模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。