用少量语言指令让机器人完成协作任务,提升操作精度与实时性。
Robotic Assistant: Completing Collaborative Tasks with Dexterous Vision-Language-Action Models
- 通过视觉骨干加入任务感知条件,增强对协作动作的理解。
- 动作后处理将关节变化压缩为4个主成分,保留96%的变异信息。
- 实测0.3秒延迟,可实时执行“拿起-传递”等长序列任务。
我们针对具身交互任务,将预训练的视觉-语言-动作模型(Open-VLA)适配于灵巧人机协作,仅需极少语言提示。方法包括:(i) 在视觉骨干中引入FiLM条件化以实现任务感知;(ii) 增加辅助意图头,预测合作者手部姿态与目标线索;(iii) 动作空间后处理,先预测紧凑的位移/旋转增量及经主成分分析降维后的手指关节,再映射为完整命令。基于多视角、远程操控的Franka和Mimic-hand数据集,并融合MediaPipe手部姿态,实验表明:增量动作行为稳定,四个主成分可解释约96%的手部关节方差。消融实验显示,动作后处理是性能主要驱动力;辅助意图有帮助,FiLM效果参差,方向性运动损失反而有害。系统在单张RTX 4090上实现约0.3秒延迟的实时组合行为,可执行“拿起”与“传递”等长时序任务。关键限制在于训练者对特定示范者的过拟合问题。
原文摘要 · Abstract (English)
We adapt a pre-trained Vision-Language-Action (VLA) model (Open-VLA) for dexterous human-robot collaboration with minimal language prompting. Our approach adds (i) FiLM conditioning to visual backbones for task-aware perception, (ii) an auxiliary intent head that predicts collaborator hand pose and target cues, and (iii) action-space post-processing that predicts compact deltas (position/rotation) and PCA-reduced finger joints before mapping to full commands. Using a multi-view, teleoperated Franka and Mimic-hand dataset augmented with MediaPipe hand poses, we demonstrate that delta actions are well-behaved and that four principal components explain ~96% of hand-joint variance. Ablations identify action post-processing as the primary performance driver; auxiliary intent helps, FiLM is mixed, and a directional motion loss is detrimental. A real-time stack (~0.3 s latency on one RTX 4090) composes "pick-up" and "pass" into a long-horizon behavior. We surface "trainer overfitting" to specific demonstrators as the key limitation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。