通过查询机制重构视觉语言表示,提升机器人导航的指令执行准确率
Action QFormer: Structured Representation Shaping under Action Supervision in Vision-Language-Action Models

- 用指令驱动的查询重排多模态信息,生成面向动作的表征
- 零样本仿真到真实导航成功率从18.8%升至56.3%
- 适合需要精准动作生成与指令对齐的具身智能研究者
在视觉-语言-动作(VLA)模型中,动作监督常被视为下游任务用于学习动作预测。本文将其视为重塑继承多模态表示的力量。我们发现这种重塑具有双重效应:虽必要于构建动作兼容表征,但若直接施加于继承的多模态路径,则可能破坏支持语言处理与物体定位的表示。为此,我们提出Action QFormer,一种基于查询的动作导向接口,利用指令条件查询将继承的多模态信息重组为面向动作的表征,再进行下游动作生成。在零样本仿真到真实导航任务中,平均闭环任务成功率由18.8%提升至56.3%,固定指令下动作生成正确率从22.5%增至75.5%,几乎消除分布外指令生成。进一步分析显示,Action QFormer改变了动作监督对继承表示的塑造方式,减少广泛上游重写,同时保留目标明确且有时具有建设性的动作监督适应。结果表明,提升VLA性能不仅需更强的预训练骨干,还需更优的方式选择与组织继承的多模态信息,并控制其在动作监督下的塑造过程。
原文摘要 · Abstract (English)
Action supervision in vision-language-action (VLA) models is often treated as a downstream objective for learning action prediction. In this paper, we study it instead as a force that shapes inherited multimodal representations. We show that this shaping has a dual effect: it is necessary for forming action-compatible representations, but when action supervision is applied too directly to the inherited multimodal pathway, it can also destabilize representations that support language-side processing and object grounding. To address this tension, we introduce Action QFormer, a query-based action-facing interface that uses instruction-conditioned queries to reorganize inherited multimodal information into action-facing representations before downstream action generation. In zero-shot sim-to-real navigation, Action QFormer improves average closed-loop task success from 18.8% to 56.3%, raises fixed-instruction action-generation correctness from 22.5% to 75.5%, and nearly eliminates out-of-distribution instruction generations. Further analyses show that Action QFormer changes how action supervision shapes inherited multimodal representations, reducing broad upstream rewriting while preserving targeted and sometimes constructive action-supervised adaptation. These results suggest that improving VLA performance requires not only stronger pretrained backbones, but also better ways of selecting and organizing inherited multimodal information while controlling how it is shaped under action supervision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。