用视觉语言模型引导学习任务相关动作,提升复杂场景下的动作建模能力
Vision-Language Models Unlock Task-Centric Latent Actions
- 利用VLM生成可提示的语义表示,分离可控动作与干扰噪声
- 在Distracting MetaWorld上使下游任务成功率提升六倍
- 发现新VLM未必更优,提示设计对效果影响显著
潜在动作模型(LAMs)已成为领先视觉-语言-动作模型预训练流程的重要组成部分。然而,当观测包含与动作相关的干扰项时,它们往往编码噪声而非有意义的潜在动作。人类仅凭简短任务描述,即可轻松区分视频中任务相关运动与无关细节。本文提出利用视觉语言模型(VLMs)的常识推理能力,生成可提示的表征,以无监督方式有效分离可控变化与噪声。我们使用这些表征作为训练目标,并评估多种主流VLMs,发现其生成可提示表征的质量及对不同提示和超参数的鲁棒性存在显著差异。有趣的是,较新的VLM可能表现不如旧模型。最后,我们证明仅要求VLM忽略干扰项即可显著提升潜在动作质量,在Distracting MetaWorld上实现最高六倍的下游任务成功率提升。
原文摘要 · Abstract (English)
Latent Action Models (LAMs) have rapidly gained traction as an important component in the pre-training pipelines of leading Vision-Language-Action models. However, they fail when observations contain action-correlated distractors, often encoding noise instead of meaningful latent actions. Humans, on the other hand, can effortlessly distinguish task-relevant motions from irrelevant details in any video given only a brief task description. In this work, we propose to utilize the common-sense reasoning abilities of Vision-Language Models (VLMs) to provide promptable representations, effectively separating controllable changes from the noise in unsupervised way. We use these representations as targets during LAM training and benchmark a wide variety of popular VLMs, revealing substantial variation in the quality of promptable representations as well as their robustness to different prompts and hyperparameters. Interestingly, we find that more recent VLMs may perform worse than older ones. Finally, we show that simply asking VLMs to ignore distractors can substantially improve latent action quality, yielding up to a six-fold increase in downstream success rates on Distracting MetaWorld.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。