arXiv:2601.18323cs.RO2026-01被引 6

用工具轨迹中间表示,让视觉生成能直接控制机器人执行动作。

TC-IDM: Grounding Video Generation for Executable Zero-shot Robot Motion

  • 以工具轨迹为中间表示,连接视觉规划与物理动作
  • 真实世界任务成功率61.11%,零样本柔性物体任务达38.46%
  • 适合需要跨任务泛化和多末端执行器的机器人控制场景

视觉-语言-动作(VLA)范式虽强大,但依赖大规模高质量机器人数据,限制泛化能力。生成式世界模型为通用具身智能提供替代路径,但其像素级计划与可执行动作间仍存在关键鸿沟。为此,我们提出工具中心逆动力学模型(TC-IDM)。通过聚焦世界模型合成视频中工具的想象轨迹,TC-IDM建立稳健的中间表示,弥合视觉规划与物理控制之间的差距。该模型通过分割与3D运动估计提取工具点云轨迹,并针对不同工具属性采用解耦动作头,将规划轨迹映射为6-自由度末端执行器运动及对应控制信号。该规划-转换范式不仅支持多种末端执行器,显著提升视角不变性,还在长时序与分布外任务中展现强泛化能力,包括与柔体物体交互。真实世界评估中,结合TC-IDM的世界模型平均成功率达61.11%,简单任务77.7%,零样本柔体任务38.46%,显著优于端到端VLA基线及其他逆动力学模型。

原文摘要 · Abstract (English)

The vision-language-action (VLA) paradigm has enabled powerful robotic control by leveraging vision-language models, but its reliance on large-scale, high-quality robot data limits its generalization. Generative world models offer a promising alternative for general-purpose embodied AI, yet a critical gap remains between their pixel-level plans and physically executable actions. To this end, we propose the Tool-Centric Inverse Dynamics Model (TC-IDM). By focusing on the tool's imagined trajectory as synthesized by the world model, TC-IDM establishes a robust intermediate representation that bridges the gap between visual planning and physical control. TC-IDM extracts the tool's point cloud trajectories via segmentation and 3D motion estimation from generated videos. Considering diverse tool attributes, our architecture employs decoupled action heads to project these planned trajectories into 6-DoF end-effector motions and corresponding control signals. This plan-and-translate paradigm not only supports a wide range of end-effectors but also significantly improves viewpoint invariance. Furthermore, it exhibits strong generalization capabilities across long-horizon and out-of-distribution tasks, including interacting with deformable objects. In real-world evaluations, the world model with TC-IDM achieves an average success rate of 61.11 percent, with 77.7 percent on simple tasks and 38.46 percent on zero-shot deformable object tasks. It substantially outperforms end-to-end VLA-style baselines and other inverse dynamics models.

机器人控制生成模型零样本动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。