通过结构设计解决视觉语言动作模型对指令改写不鲁棒的问题
Grounded Semantic Re-Binding for Robust Instruction Generalization in Vision-Language-Action Models

- 分离任务语义与视觉特征,避免联合编码引入的特征偏移
- 仅用标准示范即提升指令改写成功率44.6%,突破现有模型极限
- 新模型ParaVLA零参数量级实现近完美鲁棒性,适合轻量化部署
视觉-语言-动作(VLA)模型在机器人操作中表现优异,但当标准指令被简单改写时性能会灾难性下降。尽管通常通过昂贵的数据扩展来缓解,我们探查发现根本原因在于架构而非语义理解不足。具体而言,当前VLA能内部保留正确的任务身份,失败源于动态视觉观测与文本的联合编码引入系统性特征偏移。由于下游动作策略对这些变化极为敏感,无法将保持的语义转化为正确控制指令。为此,我们提出接地语义重绑定(GSR),通过显式融合独立提取的任务语义与原始视觉特征,从头训练全新动作专家,绕过不稳定的联合路由。该针对性干预仅使用标准示范即大幅恢复对指令改写的鲁棒性。在LIBERO-Para基准上,成功率达44.6%提升。该方法使轻量模型媲美大规模基线,并将最先进模型的PRIDE得分推至70.4的新纪录,超越近期大型预训练模型Xiaomi-Robotics-0的指令生成能力。基于此,我们还推出ParaVLA——一个0.33B参数、原生解耦的模型,在指令重述下表现出近乎完美的鲁棒性。最终证明:通过优雅的结构设计,可实现鲁棒语义定位,无需依赖低效的暴力数据扩展范式。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models excel in robotic manipulation but suffer catastrophic performance drops when canonical instructions are simply paraphrased. Although this brittleness is typically addressed through costly data scaling, our probing reveals that the root cause is architectural rather than a lack of semantic understanding. Specifically, we demonstrate that current VLAs successfully retain the correct task identity internally. The failure actually stems from the joint encoding of dynamic visual observations and text, which introduces systematic feature shifts. Because the downstream action policy is highly vulnerable to these variations, it fails to translate the preserved semantics into correct control commands. To resolve this structural bottleneck, we propose Grounded Semantic Re-binding (GSR), an elegant intervention that bypasses unstable joint routing by explicitly fusing independently extracted task semantics with native visual features to train a completely re-initialized action expert from scratch. This targeted intervention dramatically restores paraphrastic invariance using only canonical demonstrations. On the LIBERO-Para benchmark, GSR improves success rates by up to 44.6 percent. It enables lightweight models to rival massively scaled baselines and pushes state-of-the-art models to a new record PRIDE score of 70.4, outperforming the recently introduced large-scale pretrained model Xiaomi-Robotics-0 in instruction generation capabilities. Building on these insights, we also introduce ParaVLA, a natively decoupled 0.33B-parameter model exhibiting near-perfect robustness to instruction rewording. Ultimately, our work proves that robust semantic grounding can be achieved through elegant structural design, bypassing the inefficient brute-force data scaling paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。