arXiv:2512.02834cs.ROcs.AI2025-12被引 15

用测试时缩放方法提升视觉语言动作模型的推理稳定性。

Steering Vision-Language-Action Models as Anti-Exploration: A Test-Time Scaling Approach

  • 引入轻量级伪计数器验证动作片段,抑制无关动作模式。
  • 在4个仿真环境和双臂平台中成功率显著提升,稳定性强。
  • 无需梯度更新,适合扩散类模型,计算开销极低。

视觉-语言-动作(VLA)模型通过流匹配或扩散目标训练,在大规模多模态数据集(如人类遥控、脚本策略)上学习复杂行为。然而,预训练阶段包含多种数据模式,微调数据常以运动学次优或不理想方式收集,导致存在与下游任务成功动作模式无关的冗余动作模式。我们观察到,监督微调后,不同采样噪声下的推理存在关键脆弱性。本文归因于VLA策略与下游任务成功模式策略之间的分布偏移。为此提出TACO框架——一种测试时缩放(TTS)方法,使用轻量级伪计数估计器作为高保真动作块验证器。集成TACO的VLA在推理时选择伪计数最高的动作块,从而避免分布偏移,同时保持VLA泛化能力,因约束仅作用于推理阶段。该方法类似离线强化学习中的抗探索原则,且无梯度更新,相比传统强化学习更新有显著计算优势,尤其适用于难以进行强化学习更新的流或扩散基VLA。在四个仿真基准(RoboTwin2.0、Robotwin、LIBERO、SimplerEnv)及一个双臂平台上的实验表明,该方法显著提升了下游任务适配的推理稳定性与成功率。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models, trained via flow-matching or diffusion objectives, excel at learning complex behaviors from large-scale, multi-modal datasets (e.g., human teleoperation, scripted policies). However, since VLAs incorporate diverse data modes in the pre-training stage, and the finetuning dataset often contains demonstration data collected in a kinematically suboptimal or undesirable way, it exists redundant action modes that are irrelevant to the success action modes of the downstream task. Specifically, we observe a critical inference-time fragility among various sampled noises after supervised finetuning of pre-trained VLAs. In this paper, we attribute this instability to the distribution shift between the VLA policy and the policy induced by stable success modes of the downstream task dataset. Thus, we propose \textbf{TACO}, a test-time-scaling (TTS) framework that applies a lightweight pseudo-count estimator as a high-fidelity verifier of action chunks. The VLA models integrated with TACO can execute the actions with maximum pseudo-count from all sampled action chunks, thereby preventing distribution shifts while preserving the generalization ability of VLAs since the constraint is applied only during inference. Our method resembles the classical anti-exploration principle in offline reinforcement learning (RL), and being gradient-free, it incurs significant computational benefits compared to RL update, especially for flow or diffusion-based VLAs which are difficult to perform RL update due to denoising process. Extensive experiments across four simulation benchmarks (RoboTwin2.0, Robotwin, LIBERO, SimplerEnv) and a dual-arm platform demonstrate that our method significantly improves the inference stability and success rates in downstream-task adaptations.

视觉语言动作测试时缩放推理稳定抗探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。