arXiv:2605.18722cs.RO2026-05被引 7

开源首个支持双臂双手高自由度操作的视觉语言动作模型

Dexora: Open-source VLA for High-DoF Bimanual Dexterity

论文配图:Dexora: Open-source VLA for High-DoF Bimanual Dexterity
图 1 · 摘自论文原文
  • 用外骨骼+Vision Pro分层控制粗略动作和精细手指动作
  • 训练数据含10万条仿真轨迹和1万段真实操作,成功率超66%
  • 通过质量判别器筛选低质演示,提升复杂任务泛化能力

视觉-语言-动作(VLA)模型已成为具身智能的核心方向,但现有系统多局限于双机械臂或单手灵巧操作。本文提出Dexora,首个原生支持双臂双手高自由度操作的开源VLA系统。设计混合遥操作流程:通过定制外骨骼背包捕捉整体臂部运动,利用Apple Vision Pro实现无标记的手指动作追踪,并同步驱动物理平台与相同的MuJoCo数字孪生体。基于该接口构建大规模训练数据集:匹配实体的合成数据(10万条仿真轨迹,650万帧)和真实世界数据(1万段遥操作片段,292万帧)。为缓解遥操作数据噪声,提出数据质量感知训练方案:离线判别器对每段视频赋予权重,降低低质量示范的影响。实验表明,Dexora在基础与灵巧任务上均优于主流基线(灵巧任务平均成功率达66.7% vs. 51.7%),基础任务成功率达90%,并展现出强跨场景与跨实体泛化能力。消融实验证明真实数据与判别器对灵巧性提升至关重要。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently become a central direction in embodied AI, but current systems are restricted to either dual-gripper control or single-arm dexterous hand manipulation. While low-dimensional gripper control can often be handled with simpler methods, high-dimensional dexterous hand control benefits greatly from full end-to-end VLA learning. In this work, we introduce Dexora, the first open-source VLA system that natively targets dual-arm, dual-hand high-DoF manipulation. We design a hybrid teleoperation pipeline that decouples gross arm kinematics (captured with a custom exoskeleton backpack) from fine finger motion (markerless hand tracking via Apple Vision Pro), and that drives both a physical dual-arm dual-hand platform and an identical MuJoCo digital twin. Using that interface, we assemble a large training corpus: an embodiment-matched synthetic corpus (100K simulated trajectories, 6.5M frames) and a real-world dataset of 10K teleoperated episodes (2.92M frames). To mitigate noisy teleoperation demonstrations, we propose a data-quality-aware training recipe: an offline discriminator provides clip-level weights for diffusion-transformer policy training, down-weighting low-quality demonstrations. Empirically, Dexora outperforms competitive VLA baselines on both basic and dexterous benchmarks (e.g., average dexterous success 66.7% vs. 51.7%), attains 90% success on basic tasks, and shows robust out-of-distribution and cross-embodiment generalization. Ablations confirm the importance of real data and the discriminator for dexterity.

具身智能双臂操作视觉语言动作开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。