arXiv:2602.24143cs.RO2026-02被引 2

测试视觉语言动作模型在复杂场景下的泛化能力,发现其执行动作可靠但理解指令能力差。

Robust Skills, Brittle Grounding: Diagnosing Restricted Generalization in Vision-Language Action Policies via Multi-Object Picking

  • 通过逐步增加物体摆放随机性,测试模型对指令的适应能力。
  • 在完全随机布局下,动作执行成功率仍高,但按指令完成任务成功率大幅下降。
  • 建议用分层任务和分解指标来更准确评估模型的指令理解能力。

视觉语言动作(VLA)策略在少量示范下表现出色,但其性能是否源于可靠的语义-物体映射,还是依赖于特定位置关联尚不明确。我们设计了一个多物体抓取实验,逐步增加物体放置的变异性,直至全工作区随机化,并评估打破常见关联但不增加空间难度的未见物体-位置组合。在多种压力测试和数据量扩展下,发现代表性VLA模型(如SmolVLA和π₀.₅)在困难条件下,动作执行的可靠性远高于指令条件下的任务成功率,表明动作技能与指令遵循能力是解耦的。建议在操纵基准中引入任务阶梯和分解指标,分别衡量基础动作执行与指令引导的任务成功,以更好诊断指令-接地的泛化能力。

原文摘要 · Abstract (English)

Vision-language action (VLA) policies often report strong manipulation benchmark performance with relatively few demonstrations, but it remains unclear whether this reflects robust language-to-object grounding or reliance on object--location correlations that do not transfer beyond the training distribution. We present a controlled multi-object picking study that progressively increases object placement variability up to full workspace randomization and evaluates held-out object--location pairings that break familiar associations without increasing spatial difficulty. Across these stress tests and data scaling, we find that for representative VLA policies, including SmolVLA and $π_{0.5}$, execution of the manipulation primitive remains substantially more reliable than instruction-conditioned task success in harder regimes, suggesting that manipulation skill acquisition is decoupled from instruction following. We recommend augmenting manipulation benchmarks with task ladders and decomposed metrics that separately measure primitive execution and instruction-conditioned success to better diagnose instruction-grounded generalization.

视觉语言动作策略泛化能力评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。