arXiv:2602.15828cs.ROcs.CV2026-02被引 7

让机器人通过模拟学习通用抓取技能,直接用视频提示就能在真实世界操作各种物体。

Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation

  • 训练一个不依赖具体任务的3D点追踪策略,可操控任意物体到任意姿态。
  • 在数千种不同物体上训练,实现零样本迁移至真实场景,无需微调。
  • 适合需要快速部署通用灵巧操作能力的研究与工业场景。

学习能够完成大量日常任务的通用策略仍是灵巧操作领域的开放挑战。尤其是通过真实世界遥操作收集大规模操作数据成本高昂且难以扩展。虽然仿真学习提供了可行替代方案,但为每个任务设计特定环境和奖励函数同样复杂。我们提出Dex4D,一种利用仿真学习任务无关灵巧技能的框架,这些技能可灵活组合以执行多样化的现实操作任务。具体而言,Dex4D学习一个领域无关的3D点追踪策略,能将任意物体操纵至任意期望姿态。该‘任意姿态到任意姿态’策略在仿真中跨数千种不同物体及多种姿态配置下训练,覆盖广泛机器人-物体交互空间,可在测试时组合使用。部署时,该策略可零样本迁移到真实任务,仅需从生成视频中提取目标物体中心点轨迹作为提示。执行过程中,Dex4D采用在线点追踪实现闭环感知与控制。仿真与真实机器人上的大量实验表明,该方法可实现多样化灵巧操作的零样本部署,并持续优于先前基线。此外,我们展示了对新物体、场景布局、背景及轨迹的强泛化能力,凸显该框架的鲁棒性与可扩展性。

原文摘要 · Abstract (English)

Learning generalist policies capable of accomplishing a plethora of everyday tasks remains an open challenge in dexterous manipulation. In particular, collecting large-scale manipulation data via real-world teleoperation is expensive and difficult to scale. While learning in simulation provides a feasible alternative, designing multiple task-specific environments and rewards for training is similarly challenging. We propose Dex4D, a framework that instead leverages simulation for learning task-agnostic dexterous skills that can be flexibly recomposed to perform diverse real-world manipulation tasks. Specifically, Dex4D learns a domain-agnostic 3D point track conditioned policy capable of manipulating any object to any desired pose. We train this 'Anypose-to-Anypose' policy in simulation across thousands of objects with diverse pose configurations, covering a broad space of robot-object interactions that can be composed at test time. At deployment, this policy can be zero-shot transferred to real-world tasks without finetuning, simply by prompting it with desired object-centric point tracks extracted from generated videos. During execution, Dex4D uses online point tracking for closed-loop perception and control. Extensive experiments in simulation and on real robots show that our method enables zero-shot deployment for diverse dexterous manipulation tasks and yields consistent improvements over prior baselines. Furthermore, we demonstrate strong generalization to novel objects, scene layouts, backgrounds, and trajectories, highlighting the robustness and scalability of the proposed framework.

灵巧操作零样本仿真训练点追踪

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。