arXiv:2609.07747cs.RO2026-09

用人类视频学机器人灵巧操作,靠仿真补上触觉信息。

Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction

论文配图:Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction
图 1 · 摘自论文原文
  • 用仿真重建人手物交互,生成缺失的触觉数据。
  • 真实世界抓取成功率达93%,清理任务成功53%。
  • 无需实机数据,直接从视频学技能,适合想低成本学灵巧操作的人。

人类视频是灵巧操作行为的丰富来源,但缺乏对接触密集型交互至关重要的触觉信息。这引发一个根本问题:机器人能否仅从人类视频演示中学习可部署的视觉-触觉灵巧操作策略,而无需机器人端的数据采集?我们提出 DEX-X,一种通过仿真从人类视频中学习视觉-触觉灵巧操作的框架。核心思路是:仿真可作为触觉补全引擎。给定单目人类演示,DEX-X 在仿真中重建手-物交互,其中物理基础的接触动力学提供原视频中缺失的触觉监督。利用此恢复的触觉信息,我们训练视觉-触觉灵巧操作策略,并将其蒸馏为在点云观测和触觉传感下运行的可部署策略。我们在灵巧手-臂平台上演示了零样本模拟到现实的迁移,覆盖多种抓取和接触密集型工具使用任务。教师策略在仿真中六类任务平均成功率为65.9%,蒸馏后的视觉-触觉策略在真实世界立方体抓取任务中达到93%成功率,表清洗任务达53%。在未见物体几何形状的抓取任务中也观察到零样本泛化。结果表明,仿真交互是连接人类视频与可部署灵巧操作策略的关键桥梁,为从互联网规模的人类视频数据中实现可扩展的机器人技能学习提供了缺失的物理监督。

原文摘要 · Abstract (English)

Human videos are an abundant source of dexterous manipulation behaviors, but they lack tactile information that is crucial for contact-rich interaction. This raises a fundamental question: can robots learn deployable visual-tactile dexterous manipulation policies from human video demonstrations without robot-side data collection? We present DEX-X, a framework for learning visual-tactile dexterous manipulation from human videos through simulation. Our key insight is that simulation can serve as a tactile completion engine. Given monocular human demonstrations, DEX-X reconstructs hand-object interactions in simulation, where physically grounded contact dynamics provide tactile supervision unavailable in the original videos. Leveraging this recovered tactile information, we train visual-tactile dexterous manipulation policies and distill them into deployable policies operating on point-cloud observations and tactile sensing. We demonstrate zero-shot sim-to-real transfer on a dexterous hand-arm platform across diverse grasping and contact-rich tool-use tasks. The teacher policy achieves 65.9% average success across six task categories in simulation, while the distilled visual-tactile policy achieves 93% success on real-world cube picking and 53% on the challenging table-cleaning task. Zero-shot generalization to unseen object geometries is also observed on object-picking tasks. Our results suggest that simulated interaction is a key bridge between human videos and deployable dexterous manipulation policies, providing the missing physical supervision needed for scalable robot skill learning from Internet-scale human video data.

灵巧操作触觉学习仿真补全视频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。