让机器人用身体状态主动筛选视觉信息,提升决策效率与准确率。
Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies
- 将身体状态转为视觉模型可理解的词汇令牌,与指令联合筛选视觉区域。
- 仅保留约12%的视觉令牌,性能仍超完整视觉输入基线。
- 适合需要低延迟、高精度的机器人视觉-语言-动作系统应用。
视觉-语言-动作(VLA)模型通常仅在后期将本体感知作为条件信号,无法引导指令理解或视觉注意力。本文提出ThinkProprio,将本体感知离散化为视觉语言模型(VLM)词汇表中的令牌,与指令共同作用,在VLM计算前筛选视觉补丁,使模型聚焦于与动作相关的信息,提前丢弃冗余内容。系统性消融实验表明,本体感知作为被动条件信号对性能影响极小;其价值在于以令牌形式的主动查询,结合指令选择关键视觉区域。使用VLM词汇表令牌编码状态优于学习投影器,且仅保留约12%的视觉令牌即在CALVIN ABC→D任务上超越全令牌基线。在CALVIN、LIBERO及真实场景操作中,ThinkProprio均降低端到端推理延迟并提升性能。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models typically inject proprioception only as a late conditioning signal, preventing robot state from grounding instruction understanding or directing visual attention. We introduce ThinkProprio, which discretizes proprioception into VLM-vocabulary tokens and uses them jointly with the instruction to gate visual patches before VLM computation, steering the model toward action-relevant evidence while discarding redundant tokens early. We find that proprioception added as a passive conditioning signal leaves performance essentially unchanged; its value emerges when token-form state acts as an active query that, with the instruction, selects which visual patches the VLM processes. Systematic ablations show that VLM-vocabulary tokens outperform learned projectors as the state encoding, and that retaining only about \SI{12}{\percent} of the visual tokens surpasses on CALVIN ABC$\to$D. Across CALVIN, LIBERO, and real-world manipulation, ThinkProprio reduces end-to-end inference latency while improving the matched full-token baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。