arXiv:2511.00139cs.ROcs.AI2025-11被引 11

人机协同操控机器人完成精细抓握,效率更高且成功率超90%。

End-to-End Dexterous Arm-Hand VLA Policies via Shared Autonomy: VR Teleoperation Augmented by Autonomous Hand VLA Policy for Efficient Data Collection

  • 人类用VR控制手臂大动作,机器自主处理手部微调。
  • 收集高质量数据仅需少量人力,抓握成功率90%以上。
  • 适合需要高精度操作的机器人研发与教学场景。

实现类人级灵巧操作仍是通用机器人面临的主要挑战。尽管视觉-语言-动作(VLA)模型在从示范中学习技能方面展现出潜力,但其可扩展性受限于高质量训练数据稀缺。现有数据采集方法存在固有局限:手动遥操作加重人力负担,自动化规划常产生不自然动作。本文提出共享自主框架,将控制分为宏观与微观运动。人类操作员通过直观的虚拟现实(VR)遥操作引导机器人手臂姿态,而自主的DexGrasp-VLA策略则利用实时触觉与视觉反馈处理精细的手部控制。该分工显著降低认知负荷,实现高效采集高质量协同臂手示范数据。基于此数据,我们训练了端到端的VLA策略,并引入新型臂手特征增强模块,捕捉宏观与微观运动的独立与共享表征,实现更自然的协调。我们的纠正式遥操作系统支持通过人机协同故障恢复持续优化策略。实验表明,该框架以极低人力投入生成高质量数据,在多种物体(包括未见过的实例)上均达到90%成功率。全面评估验证了系统在发展灵巧操作能力方面的有效性。

原文摘要 · Abstract (English)

Achieving human-like dexterous manipulation remains a major challenge for general-purpose robots. While Vision-Language-Action (VLA) models show potential in learning skills from demonstrations, their scalability is limited by scarce high-quality training data. Existing data collection methods face inherent constraints: manual teleoperation overloads human operators, while automated planning often produces unnatural motions. We propose a Shared Autonomy framework that divides control between macro and micro motions. A human operator guides the robot's arm pose through intuitive VR teleoperation, while an autonomous DexGrasp-VLA policy handles fine-grained hand control using real-time tactile and visual feedback. This division significantly reduces cognitive load and enables efficient collection of high-quality coordinated arm-hand demonstrations. Using this data, we train an end-to-end VLA policy enhanced with our novel Arm-Hand Feature Enhancement module, which captures both distinct and shared representations of macro and micro movements for more natural coordination. Our Corrective Teleoperation system enables continuous policy improvement through human-in-the-loop failure recovery. Experiments demonstrate that our framework generates high-quality data with minimal manpower and achieves a 90% success rate across diverse objects, including unseen instances. Comprehensive evaluations validate the system's effectiveness in developing dexterous manipulation capabilities.

灵巧操作人机协同数据采集视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。