arXiv:2603.08122cs.RO2026-03被引 7

用强化学习辅助遥操作,实现更灵活的双手灵巧抓取。

Towards Human-Like Manipulation through RL-Augmented Teleoperation and Mixture-of-Dexterous-Experts VLA

  • 引入强化学习训练的原子技能,简化数据采集并作为底层执行单元。
  • 在复杂接触任务中,成功率比基线提升一倍。
  • 融合力觉与触觉模态,不破坏预训练模型知识。

尽管视觉-语言-动作(VLA)模型在机器人操作中表现优异,但其应用仍局限于低自由度末端执行器完成简单的视觉引导抓放任务。将这些模型扩展至类人、双臂灵巧操作——特别是富含接触的手中操作——面临高保真数据获取、多技能学习和多模态传感融合等关键挑战。本文提出一个集成框架,包含两个核心组件:首先,提出IMCopilot(手中操作协作者),一套通过强化学习训练的原子技能,兼具共享自治辅助遥操作数据收集和作为VLA可调用底层执行原语的双重功能;其次,提出MoDE-VLA(灵巧专家混合型VLA),通过残差注入机制无缝整合异构的力觉与触觉模态至预训练VLA骨干网络,实现接触感知优化而不损害模型预训练知识。我们在四个复杂度递增的任务上验证该方法,在灵巧接触任务中成功率较基线提升一倍。

原文摘要 · Abstract (English)

While Vision-Language-Action (VLA) models have demonstrated remarkable success in robotic manipulation, their application has largely been confined to low-degree-of-freedom end-effectors performing simple, vision-guided pick-and-place tasks. Extending these models to human-like, bimanual dexterous manipulation-specifically contact-rich in-hand operations-introduces critical challenges in high-fidelity data acquisition, multi-skill learning, and multimodal sensory fusion. In this paper, we propose an integrated framework to address these bottlenecks, built upon two components. First, we introduce IMCopilot (In-hand Manipulation Copilot), a suite of reinforcement learning-trained atomic skills that plays a dual role: it acts as a shared-autonomy assistant to simplify teleoperation data collection, and it serves as a callable low-level execution primitive for the VLA. Second, we present MoDE-VLA (Mixture-of-Dexterous-Experts VLA), an architecture that seamlessly integrates heterogeneous force and tactile modalities into a pretrained VLA backbone. By utilizing a residual injection mechanism, MoDE-VLA enables contact-aware refinement without degrading the model's pretrained knowledge. We validate our approach on four tasks of escalating complexity, demonstrating doubled success rate improvement over the baseline in dexterous contact-rich tasks.

灵巧操作强化学习多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。