让人类实时干预机械手操作,显著提升机器人灵巧抓取的性能。
DexHiL: A Human-in-the-Loop Framework for Vision-Language-Action Model Post-Training in Dexterous Manipulation
- 引入干预感知的数据采样策略,优先学习人类纠正段落。
- 真实机器人测试中成功率平均提升25%,超越纯离线微调基线。
- 首个集成手臂与灵巧手的人机协同后训练框架,适合复杂抓取任务。
尽管视觉-语言-动作(VLA)模型在机器人操作中展现出良好泛化能力,但在具体复杂下游任务中的部署仍需有效的后训练。与此同时,人机协同(HiL)学习已被证明是优化机器人策略的强大手段。然而,将这一范式扩展至灵巧操作仍面临挑战:多指控制维度高、接触密集,且执行分布与常规臂部运动差异显著,导致现有灵巧型VLA系统在可靠性与适应性上受限。本文提出DexHiL,首个集成臂-手人机协同的灵巧型VLA后训练框架,实现对臂部与灵巧手的统一协同干预。DexHiL引入一种干预感知的数据采样策略,优先选择人类纠正片段用于后训练,并配备轻量级遥操作界面,支持执行过程中的即时人工修正。真实机器人实验表明,DexHiL作为后训练框架效果显著,在多个任务上平均成功率比标准离线微调基线提升25%。
原文摘要 · Abstract (English)
While Vision-Language-Action (VLA) models have demonstrated promising generalization capabilities in robotic manipulation, deploying them on specific and complex downstream tasks still demands effective post-training. In parallel, Human-in-the-Loop (HiL) learning has proven to be a powerful mechanism for refining robot policies. However, extending this paradigm to dexterous manipulation remains challenging: multi-finger control is high-dimensional, contact-intensive, and exhibits execution distributions that differ markedly from standard arm motions, leaving existing dexterous VLA systems limited in reliability and adaptability. We present DexHiL, the first integrated arm-hand human-in-the-loop framework for dexterous VLA models, enabling coordinated interventions over the arm and the dexterous hand within a single system. DexHiL introduces an intervention-aware data sampling strategy that prioritizes corrective segments for post-training, alongside a lightweight teleoperation interface that supports instantaneous human corrections during execution. Real-robot experiments demonstrate that DexHiL serves as an effective post-training framework, yielding a substantial performance leap, outperforming standard offline-only fine-tuning baselines by an average of 25% in success rates across distinct tasks. Project page: https://chenzhongxi-sjtu.github.io/dexhil/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。