arXiv:2604.23620cs.RO2026-04被引 1

将机器人操作分为移动与接触两阶段,提升精准抓取成功率。

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

论文配图:Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation
图 1 · 摘自论文原文
  • 分阶段控制:先粗略移动,再精细操作,结构更清晰。
  • 成功率68.9%,比单一策略高24%,训练更快更省数据。
  • 适合需要高精度动作的工业机器人场景。

我们提出一种视觉语言动作框架Move-Then-Operate,将机器人操作显式分解为粗略重定位(move)和接触关键交互(operate)两个阶段。不同于将两种行为混为一体的单体策略,该方法采用双专家政策,由可学习的阶段选择器路由,引入结构归纳偏置以分离各阶段动态。阶段标签通过基于多模态大模型的管道自动生成,依赖末端执行器速度等轻量上下文线索,确保与人类运动模式对齐。在RoboTwin2基准上,该方法平均成功率达68.9%,较单体基线π₀提升24%。其性能媲美甚至超过在10倍更多数据上训练的模型,且仅需40%的训练步数达到峰值表现,证明显式解耦移动与操作阶段是掌握高精度操作的有效高效策略。

原文摘要 · Abstract (English)

We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). Unlike monolithic policies that conflate these heterogeneous regimes, our architecture employs a dual-expert policy routed by a learnable phase selector, introducing a structural inductive bias that isolates phase-specific dynamics. Phase labels are automatically generated via an MLLM-based pipeline conditioned on lightweight contextual cues such as end-effector velocity and subtask decomposition to ensure alignment with human motor patterns. Evaluated on the RoboTwin2 benchmark, our method achieves an average success rate of $68.9\%$, outperforming the monolithic $π_0$ baseline by $24\%$. It matches or exceeds models trained on $10\times$ more data and reaches peak performance in $40\%$ fewer training steps, demonstrating that architectural disentanglement of move and operate phases is a highly effective and efficient strategy for mastering high-precision manipulation.

机器人操作分阶段控制视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。