arXiv:2608.28108cs.ROcs.CV2026-08

让机器人同时理解语言、看图和手势指令,统一操作界面。

DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

论文配图:DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
图 1 · 摘自论文原文
  • 将三种指令统一为文本提示+指向性掩码,用同一模型处理。
  • 在真实场景中,新物体和外观变化下成功率提升至100%。
  • 适合需要灵活交互的机器人操作任务,如家庭服务或工业装配。

视觉-语言-动作模型(VLAs)允许用户通过自然语言指定操作任务,但在相同类别或相似外观的物体中区分目标或放置位置时,需详细表达,而现有VLA可能无法可靠使用。我们提出DeicticVLA,通过文本补全和指向性手势定位,将语言指令(LI)、视觉-语言指令(VLI)和视觉指令(VI)统一为文本提示与指向性掩码,使单一预训练VLA可处理三种指令模式。采用共享主干网络、演示数据及匹配训练步骤,在仿真环境中比较两种RGB视觉提示方法、两种分通道掩码提示方法和三种训练策略。两阶段训练下,四种提示方法均实现高分布内成功率,但在未见布局中使用指向性掩码的能力存在差异。训练策略消融显示,两阶段训练有助于提升该能力,保留第二阶段语言数据可缓解遗忘,且不降低VLI与VI性能。在三个真实任务中,一个策略支持所有模式;当表达未见、外观变化或新物体时,VLI与VI表现优于LI,对未见类别成功率均为100%,而联合训练的LI仅达16.7%。结果验证了三模式统一接口的有效性,并指导DeicticVLA设计。

原文摘要 · Abstract (English)

Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably. We propose DeicticVLA, which canonicalizes Language Instruction (LI), Vision-Language Instruction (VLI), and Visual Instruction (VI) into a text prompt and deictic masks through text-prompt completion and deictic gesture grounding, enabling a single pretrained VLA to handle all three instruction modes. With a shared backbone, demonstrations, and matched training steps, we compare two RGB visual prompting methods, two separate-channel mask prompting methods, and three training strategies in simulation. Under two-stage training, the four prompting methods achieve high in-distribution success but differ in their ability to use deictic masks in unseen layouts. Across methods, training-strategy ablations show that two-stage training improves such use, while retaining second-stage LI data mitigates forgetting without reducing VLI and VI performance. In three real-world tasks, one policy supports all modes. VLI and VI outperform LI under unseen expressions, appearance changes, and novel objects. For unseen categories, both achieve 100% success, compared with 16.7% for jointly trained LI. These results demonstrate the unified three-mode interface and guide DeicticVLA design.

多模态指令机器人控制视觉语言动作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。