通过测试时验证提升视觉语言动作对齐,效果优于扩大策略训练。
Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment
- 测试时联合扩增重述指令与生成动作,提升样本多样性。
- 在SIMPLER上比扩展训练提升22%(分布内)和13%(分布外)。
- 适合追求高效部署的机器人系统研究者。
通用机器人的长期愿景依赖于其理解并执行自然语言指令的能力。视觉-语言-动作(VLA)模型在此目标上取得显著进展,但生成的动作仍可能与指令不符。本文研究测试时验证以缩小‘意图-动作’差距。我们首先刻画了具身指令遵循的测试时缩放规律,发现联合扩增重述指令数量与生成动作数能大幅提高测试时样本多样性,通常比独立扩增更高效地恢复正确动作。为利用这些缩放规律,我们提出CoVer——一种用于视觉-语言-动作对齐的对比验证器,并证明其架构可随计算资源与数据量增长而良好扩展。随后引入CoVer-VLA,一个基于训练好的验证器的分层测试时验证流水线。部署时,框架从视觉-语言模型(VLM)预生成多样重述指令,为每条指令反复生成动作候选,再用验证器筛选最优高层提示与低层动作片段。相比在同一数据上扩展策略预训练,我们的验证方法在SIMPLER基准上实现22%(分布内)和13%(分布外)的提升,真实世界实验中进一步提升45%。在PolaRiS基准上,任务进展提升14%,成功率提升9%。
原文摘要 · Abstract (English)
The long-standing vision of general-purpose robots hinges on their ability to understand and act upon natural language instructions. Vision-Language-Action (VLA) models have made remarkable progress toward this goal, yet their generated actions can still misalign with the given instructions. In this paper, we investigate test-time verification as a means to shrink the "intention-action gap." We first characterize the test-time scaling laws for embodied instruction following and demonstrate that jointly scaling the number of rephrased instructions and generated actions greatly increases test-time sample diversity, often recovering correct actions more efficiently than scaling each dimension independently. To capitalize on these scaling laws, we present CoVer, a contrastive verifier for vision-language-action alignment, and show that our architecture scales gracefully with additional computational resources and data. We then introduce CoVer-VLA, a hierarchical test-time verification pipeline using the trained verifier. At deployment, our framework precomputes a diverse set of rephrased instructions from a Vision-Language-Model (VLM), repeatedly generates action candidates for each instruction, and then uses the verifier to select the optimal high-level prompt and low-level action chunks. Compared to scaling policy pre-training on the same data, our verification approach yields 22% gains in-distribution and 13% out-of-distribution on the SIMPLER benchmark, with a further 45% improvement in real-world experiments. On the PolaRiS benchmark, CoVer-VLA achieves 14% gains in task progress and 9% in success rate.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。