arXiv:2607.26315cs.RO2026-07

让机器人通过调节动作模式实现灵活操作,跨任务复用。

MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization

论文配图:MoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
图 1 · 摘自论文原文
  • 用时空动作分词器提取动作模式,输入条件控制执行方式。
  • 六种真实任务中,不同模式下速度、加速度等行为可区分且稳定。
  • 未见过的动作模式也能成功迁移,适合需要灵活执行的机器人场景。

为在多样环境中有效作业,机器人不仅需精准完成操作,还应根据任务、物体和交互情境调整动作展开方式。本文探讨这种执行层面的差异能否作为可复用的行为因子跨任务学习。提出名为MoMo的两阶段模仿学习框架,包含时空动作分词器与行为克隆变压器,以任务和连续动作模式为输入。在六种真实机器人操作任务中,改变该条件可生成稳定、动态及中间态行为,人类评分员可区分其差异,且在关节速度、加速度与末端执行器接近角度上显著不同。对于仅演示过一种模式的任务,MoMo仍能成功转移未见请求模式,同时保持任务成功率。结果表明,动作模式可在不同任务间复用,支持对未见任务-模式组合的组合泛化,验证了动作模式作为可复用执行控制因子的有效性。

原文摘要 · Abstract (English)

To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execution-level variation can be learned as a reusable behavioral factor shared across tasks. We present \textbf{MoMo}, a two-stage imitation-learning framework consisting of a spatiotemporal action tokenizer and a behavior-cloning transformer that takes task and a continuous motion-mode condition as inputs. Across six real-robot manipulation tasks, varying this condition produces steady, dynamic, and intermediate behaviors that human raters can distinguish and that differ in joint speed, acceleration, and end-effector approach pitch. On tasks demonstrated in only one mode, MoMo transfers the unseen requested mode while largely preserving task success. Together, these results provide evidence of compositional generalization to unseen task--mode combinations and show that motion mode can be reused across tasks to control how a manipulation skill is performed.

机器人操作动作模式模仿学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。