arXiv:2606.13435cs.RO2026-06

让机器人通过手势理解人类指令,提升交互准确率。

GIVE: Grounding Human Gestures in Vision-Language-Action Models

论文配图:GIVE: Grounding Human Gestures in Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用视觉与语义双路径融合手势信息,增强指令理解。
  • 实测目标识别准确率提升40%,任务成功率提高80%。
  • 适合需要自然人机交互的机器人应用,尤其在复杂场景中。

人类交流本质是多模态的,语言常伴随手势等非语言线索以传达意图。然而,当前视觉-语言-动作(VLA)模型将机器人操作视为纯文本驱动任务,忽视了手势在人机交互(HRI)中的关键作用,导致在语言指令模糊或不完整时意图定位不准、操作不可靠。为此,我们提出GIVE(Gesture Intent via Visual-Semantic Enhancement),一种无需修改架构即可增强预训练VLA模型手势理解能力的有效方法。GIVE通过两条互补路径融入手势信息:视觉路径在机器人观测上叠加手部骨骼和指尖射线,实现物体的显式定位;语义路径生成手势与任务指令的高层描述,提升意图定位鲁棒性。联合利用视觉与语义引导,使VLA策略能更好关联手势与操作行为,并适应动态交互意图。真实世界HRI实验表明,GIVE显著优于基线,在目标物体识别准确率上提升40%,整体任务成功率提高80%,且对未见过的空间布局和不同参与者表现出强鲁棒性与泛化能力。

原文摘要 · Abstract (English)

Human communication is inherently multimodal, where language is often accompanied by non-verbal cues such as gestures to convey intentions. However, current Vision-Language-Action (VLA) models treat robotic manipulation as a pure text-driven task, overlooking the important role of gestures in Human-Robot Interaction (HRI). This often leads to inaccurate intent grounding and unreliable manipulation when language instructions are ambiguous or underspecified. To address this challenge, we propose GIVE (Gesture Intent via Visual-Semantic Enhancement), an effective approach that enhances pre-trained VLA models with human gesture understanding without architectural modifications. Specifically, GIVE incorporates gesture information through two complementary pathways: a visual pathway that overlays hand skeletons and fingertip rays onto robot observations for explicit object grounding, and a semantic pathway that generates high-level descriptions of human gestures and task instructions for robust intent grounding. By jointly leveraging visual and semantic guidance, GIVE enables VLA policies to better associate gestures with manipulation behaviors and adapt to dynamic interaction intents. In real-world HRI experiments, GIVE substantially outperforms the baseline, improving target object recognition accuracy by 40% and overall task success rate by 80%, while demonstrating strong robustness and generalization to unseen spatial layouts and diverse participants.

人机交互手势理解多模态机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。