arXiv:2605.28527cs.RO2026-05

冻结的视觉语言动作模型暗藏成功预测能力,无需重训练即可提升机器人决策。

What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies

论文配图:What Frozen VLAs Already Know About Success: A Probing Study of Value-Like Structure in Foundation Robot Policies
图 1 · 摘自论文原文
  • 用轻量线性探测器从冻结特征中提取成功预测信号。
  • 在推杆和开瓶任务上,成功率从26.7%提升至44.3%。
  • 适用于需要高精度决策的机器人控制场景。

视觉-语言-动作(VLA)策略通过模仿动作进行训练,其损失函数从未要求估计奖励、进展或未来成功。然而,这些冻结表示中仍包含此类信息,可被读取并用于指导动作选择而无需重新训练。基于LIBERO-Goal上的混合成功与失败操作轨迹,我们利用轻量级线性探测器从冻结特征中恢复蒙特卡洛结果目标。该目标在OpenVLA、Pi0.5、DINOv2和CLIP特征中均能稳定预测,而在基于进展、剩余时间、任务身份或本体感知构建的基线中则显著较弱。为排除任务和时间捷径,我们在同任务、同时间步匹配条件下评估探测器:Pi0.5探测器仍达到约92%的成对排序准确率,而标签打乱对照组保持随机水平。将同一探测器作为测试时动作前缀选择器使用,使离线发现转化为实际行为:在推杆任务上,成功率从贪婪解码的26.7%提升至44.3%,开瓶任务也出现第二次正向提升。增益并非普遍适用,且需额外推理计算,但核心发现清晰:冻结的VLA已编码其模仿目标未明确要求的成功信息。

原文摘要 · Abstract (English)

Vision--language--action (VLA) policies are trained to imitate actions; their loss never asks them to estimate reward, progress, or future success. Their frozen representations nevertheless carry such information, and it can be read out and used to guide action choice without retraining the policy. From mixed successful and failed manipulation trajectories on LIBERO-Goal, we recover Monte-Carlo outcome targets using lightweight linear probes on frozen features. The targets are consistently predictable from OpenVLA, Pi0.5, DINOv2, and CLIP features, and substantially less so from baselines built on progress, time-to-go, task identity, or proprioception. To rule out task and temporal shortcuts, we evaluate the probes under same-task, same-timestep matched comparisons: Pi0.5 probes still reach roughly 92% pairwise ordering accuracy, while label-shuffled controls stay at chance. Used as a test-time selector over sampled Pi0.5 action prefixes, the same probe turns this offline finding into behavior: on push-plate, success rises from 26.7% under greedy decoding to 44.3%, with a second positive case on wine-rack. The gains are not universal and require additional inference compute, but the underlying finding is clean: frozen VLAs already encode information about success that their imitation objective never explicitly demands.

机器人成功预测冻结模型视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。