arXiv:2606.12109cs.ROcs.AI2026-06被引 1

让视觉语言动作模型学会用灵活双手抓取物体,提升操作成功率。

InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation

论文配图:InDex: Empowering VLA Models with Intent-Conditioned Arm-Hand Coordination for Dexterous Manipulation
图 1 · 摘自论文原文
  • 分离抓取时机与手部动作规划,用意图信号协调机械臂与手的配合。
  • 在4个仿真任务和真实机器人平台上,抓取成功率显著提升。
  • 适合研究灵巧操作、具身智能或视觉语言动作模型的开发者。

预训练的视觉-语言-动作(VLA)模型提供了有用的语义和空间先验,但其并行夹爪的动作接口未说明如何通过灵巧手实现这些先验。直接附加手指关节会混淆两个结构不同的决策:何时建立接触,以及如何以形态特定的手部轨迹建立接触。我们提出InDex,一种意图条件化的适配框架,分离这两个决策而不丢弃完整的手部监督。InDex从重定向演示中推导出归一化的抓取意图。第一阶段预测同步的末端执行器-意图片段;基于这些预测、VLA上下文和本体感知,扩散解码器生成多关节手部动作。因此,标量意图是时间协调接口,而非压缩的手部姿态。在四个仿真任务、三个VLA主干网络和一个物理机械臂-手平台上,InDex保持了VLA的接近能力,同时显著提升了从接近到稳定抓取及任务完成的转化率。消融实验表明:意图对齐接触过渡,而扩散模型表示与同一任务空间规划兼容的多种手部轨迹。结果表明,后接近阶段的臂-手协调,而非仅物体定位,是将并行夹爪型VLA应用于灵巧操作的主要瓶颈。

原文摘要 · Abstract (English)

Pre-trained Vision-Language-Action (VLA) models provide useful semantic and spatial priors, yet their parallel-gripper action interfaces do not specify how those priors should be realized by a dexterous hand. Directly appending finger joints conflates two decisions with different structure: when contact should be established and how a morphology-specific hand trajectory should establish it. We introduce InDex, an intent-conditioned adaptation framework that separates these decisions without discarding full hand supervision. InDex derives a normalized grasp intent from retargeted demonstrations. A first stage predicts synchronized end-effector--intent chunks; conditioned on these predictions, VLA context, and proprioception, a diffusion decoder generates multi-joint hand actions. The scalar intent is therefore a temporal coordination interface rather than a compressed hand pose. Across four simulated tasks, three VLA backbones, and a physical arm--hand platform, InDex preserves the VLA's reaching competence while markedly improving conversion from approach to stable grasp and task completion. Ablations isolate complementary roles: intent aligns the contact transition, whereas diffusion represents the multiple hand trajectories compatible with the same task-space plan. These results identify post-reach arm--hand coordination, rather than object localization alone, as the principal bottleneck in adapting parallel-gripper VLAs to dexterous manipulation.

灵巧操作视觉语言动作扩散模型机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。