arXiv:2605.22812cs.ROcs.CV2026-05被引 2

用手势增强机器人视觉语言模型,让机械臂更准地理解复杂指令。

GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations

论文配图:GesVLA: Gesture-Aware Vision-Language-Action Model Embedded Representations
图 1 · 摘自论文原文
  • 将手势特征直接嵌入潜空间,与视觉语言联合推理
  • 在杂乱场景中目标定位准确率显著提升,交互效率提高30%以上
  • 适合需要手势交互的智能机器人研发人员参考

视觉-语言-动作(VLA)模型通过融合感知与动作,在通用机器人操作中展现出巨大潜力。然而,现有系统主要依赖文本指令,在存在多个相似物体的复杂场景中难以解决空间歧义问题。为此,本文引入手势作为并行指令模态,提出手势感知型视觉-语言-动作模型(GesVLA)。该方法将手势特征直接编码至潜在空间,使其参与高层推理与底层动作生成,并采用双视觉语言模型架构实现手势表征与动作策略的紧密耦合。在数据层面,构建可扩展的手势数据生成流水线,通过将手部模型渲染到真实场景图像上,降低仿真到现实的视觉差异,同时生成具有多样化运动模式和对应指向标注的丰富数据。此外,采用两阶段训练策略,使模型具备手势感知与动作预测双重能力。在多个真实机器人任务上评估,包括控制块体操作任务及产品、生鲜选品等实际场景。实验结果表明,引入手势能持续提升目标定位准确率与人机交互效率,尤其在复杂杂乱环境中表现突出。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial ambiguity in complex scenes with multiple similar objects. To address this limitation, we introduce gesture as a parallel instruction modality and propose a Gesture-aware Vision-Language-Action model (GesVLA). Our approach encodes gesture features directly into the latent space, enabling them to participate in both high-level reasoning and low-level action generation, and adopts a dual-VLM architecture to achieve tight coupling between gesture representations and action policies. At the data level, we construct a scalable gesture data generation pipeline by rendering hand models onto real-world scene images. This reduces the sim-to-real visual gap while producing rich data with diverse motion patterns and corresponding pointing annotations. In addition, we employ a two-stage training strategy to equip the model with both gesture perception and action prediction capabilities. We evaluate our approach on multiple real-world robotic tasks, including a controlled block manipulation task for validation and more practical scenarios such as product and produce selection. Experimental results show that incorporating gesture consistently improves target grounding accuracy and human-robot interaction efficiency, especially in complex and cluttered environments. Project page: https://gwxuan.github.io/GesVLA/.

机器人手势识别多模态具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。