通过贝叶斯分解提升视觉语言动作模型的泛化能力
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
- 将策略分解为视觉-动作先验与语言条件似然,实现看图执行与提示细化
- 在未见指令、物体和环境上表现优于现有方法,泛化能力显著提升
- 适用于需要强指令遵循与跨场景适应的机器人控制任务
视觉语言动作(VLA)模型在分布外泛化方面常受困于微调过程中视觉语言模型(VLM)主干的灾难性遗忘。尽管联合训练外部推理数据有所帮助,但需经验调参且存在数据开销。除外部依赖外,我们发现VLA数据集中的内在问题:模态失衡——语言多样性远低于视觉与动作多样性。这种失衡导致模型偏向视觉捷径并遗忘语言信息。为此,我们提出BayesVLA,一种贝叶斯分解方法,将策略分解为视觉-动作先验(支持看图执行)和语言条件似然(实现提示细化),天然保留泛化能力并促进指令遵循。我们进一步引入接触前/后阶段以更好利用预训练基础模型。信息论分析形式化验证了其缓解捷径学习的有效性。大量实验表明,相比现有方法,该模型在未见指令、物体和环境上具备更优泛化性能。
原文摘要 · Abstract (English)
The pursuit of out-of-distribution generalization in Vision-Language-Action (VLA) models is often hindered by catastrophic forgetting of the Vision-Language Model (VLM) backbone during fine-tuning. While co-training with external reasoning data helps, it requires experienced tuning and data-related overhead. Beyond such external dependencies, we identify an intrinsic cause within VLA datasets: modality imbalance, where language diversity is much lower than visual and action diversity. This imbalance biases the model toward visual shortcuts and language forgetting. To address this, we introduce BayesVLA, a Bayesian factorization that decomposes the policy into a visual-action prior, supporting seeing-to-act, and a language-conditioned likelihood, enabling prompt-to-specify. This inherently preserves generalization and promotes instruction following. We further incorporate pre- and post-contact phases to better leverage pre-trained foundation models. Information-theoretic analysis formally validates our effectiveness in mitigating shortcut learning. Extensive experiments show superior generalization to unseen instructions, objects, and environments compared to existing methods. Project page is available at: https://xukechun.github.io/papers/BayesVLA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。