探究机器人感知状态如何影响视觉语言动作模型性能
How Should Vision-Language-Action Models Use Proprioceptive State?

- 设计五种状态输入方式,在相同条件下对比效果
- 发现历史状态信息提升性能,96帧历史优于单帧
- 适合研究机器人多模态控制与模型设计的学者
近期视觉-语言-动作(VLA)模型普遍将机器人本体感知状态作为输入,但接入方式不一——或序列化为文本提示、投影至视觉-语言前缀、或直接输入动作专家模块,且几乎都仅使用单一当前帧。三个核心问题仍待解答:(1)当前状态是否真能提升闭环控制性能,适用于哪些任务?(2)历史状态信息需要多少才有效,其收益是源于真实时间变化还是额外条件容量?(3)状态应接入模型的哪个部分——视觉-语言主干还是动作生成模块?我们通过在流匹配型VLA上进行受控实验,固定主干网络、训练数据、动作表示与评估协议,实现五种代表性接口(离散状态提示、VLM前缀、动作前缀、状态专家、特征调制),并在45个基础任务(涵盖三类任务家族)及20个复合任务上评估。同时,将状态历史长度从1至96帧系统性地扫描,分析历史信息对模型表现的影响。实验结果系统回答了上述所有问题,并提炼出可验证的状态感知型VLA设计原则。
原文摘要 · Abstract (English)
Recent Vision-Language-Action (VLA) models almost universally take robot proprioceptive state as input, yet wire it in incompatible ways -- serialized into text prompts, projected into the vision-language prefix, or fed directly to the action expert -- and almost always as a single current frame. Three questions remain open: (1) whether, and on which tasks, current state actually improves closed-loop control; (2) how much state history helps, and whether its benefit reflects genuine temporal variation rather than added conditioning capacity; and (3) where state should enter the model -- the vision-language backbone or the action-generation module. We answer these questions through controlled experiments on a flow-matching VLA, fixing the backbone, training data, action representation, and evaluation protocol throughout. We implement five representative interfaces -- discrete state prompt, VLM prefix, action prefix, state expert, and feature modulation -- under matched implementation details, and evaluate them on 45 atomic tasks spanning three task families plus 20 composite tasks; we then sweep the state-history length from 1 to 96 frames to examine how historical state information affects model performance. The experiments yield systematic answers to all three questions, distilled into testable design principles for state-aware VLAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。