arXiv:2608.05369cs.ROcs.CV2026-08

让机器人预测手腕动作未来,提升精细操作能力

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

论文配图:World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation
图 1 · 摘自论文原文
  • 用任务条件化潜变量接口连接视觉语言与手腕预测
  • 在多个数据集上实现80赫兹以上动作生成,提升接触敏感操作精度
  • 适合需要精细手腕控制的机器人任务,如穿珠子、拧螺丝

视觉-语言-动作(VLA)模型通常将主视角和腕部视角视为并行输入,忽略了它们在机器人操作中的不同作用。精细操作则依赖于在全局任务背景下预判腕部局部交互的演化。为此,我们提出世界到手腕的VLA模型(W2-VLA),支持细粒度机器人操作中的任务条件化未来腕部建模。给定多视角观测和任务指令,W2-VLA将一组潜变量建模令牌作为视觉语言模型与腕部预测器之间的紧凑接口。该接口结合已观测腕部历史,预测未来腕部潜变量,并将其转化为用于动作预测的未来感知上下文。此外,我们引入了W2-CoT合成管道,生成包含操作进展、物理变化线索和腕部局部证据的结构化标注。这些标注提供辅助监督,塑造任务条件化的潜变量接口。在LIBERO、RoboTwin 2.0及真实世界操作任务上的实验表明,该方法在单臂与双臂设置下均提升了细粒度与接触敏感操作性能,同时保持动作生成速率高于80赫兹。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wrist-local interactions may evolve under the global task context. To address this limitation, we present World-to-Wrist VLA (W2-VLA), a VLA model for fine-grained robot manipulation with task-conditioned future wrist modeling. Given current multi-view observations and a task instruction, W2-VLA contextualizes a set of latent modeling tokens as a compact interface between the vision-language model and the wrist predictor. Conditioned on this interface and the observed wrist history, the predictor forecasts future wrist latents, which are transformed into future-aware context for action prediction. In addition, we introduce W2-CoT, a synthesis pipeline that produces structured annotations describing manipulation progress, physical transition cues, and wrist-local evidence. These annotations provide auxiliary supervision that shapes the task-conditioned latent interface. Experiments on LIBERO, RoboTwin 2.0, and real-world manipulation tasks demonstrate improved fine-grained and contact-sensitive manipulation across both single-arm and bimanual settings, while maintaining action-generation rates above 80 Hz.

机器人操作多视图建模动作预测细粒度控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。