预判动作好坏,让视觉语言模型执行更可靠。
Pre-VLA: Preemptive Runtime Verification for Reliable Vision-Language-Action and World-Model Rollouts

- 用多模态模型提前评估动作是否安全有效。
- 在LIBERO上成功率提升至37.62%,平均验证耗时183.9毫秒。
- 适合需要高可靠性的机器人任务与世界模型推演场景。
尽管大型视觉-语言-动作(VLA)模型和生成式世界模型(WM)推动了长时程具身智能的发展,但其实际部署仍受限于学习生成动作的不确定性。低质量动作可能导致物理执行失败或引发冗余渲染开销的误导性世界模型推演。为此,我们提出Pre-VLA,一种统一的运行时验证架构,可在物理执行或世界模型想象前对候选动作块进行预先有效性评估。Pre-VLA采用高效的多模态骨干网络与模态感知池化,并结合轻量级双分支头,预测动作的安全置信度与批判性优势分数。为应对严重类别不平衡与不稳定边界决策问题,训练中引入融合焦点分类、优势回归与软阈值校准的多任务目标。部署时,双模式预采样调度器在有限计算预算下过滤低质量动作并触发自适应重采样。在LIBERO基准测试中,Pre-VLA将四个任务套件的平均闭环成功率从30.79%提升至37.62%,减少任务执行步数,实现每动作块183.9毫秒的平均前向验证时间,并有效缓解世界模型推演中的误差累积。
原文摘要 · Abstract (English)
While large vision-language-action (VLA) models and generative world models (WM) have advanced long-horizon embodied intelligence, their practical deployment remains challenged by uncertainty in learning-based action generation. Low-quality actions may cause physical failures during execution or lead to misleading world-model rollouts with redundant rendering costs. To address this issue, we propose Pre-VLA, a unified runtime verification architecture that performs preemptive action validity assessment before physical execution or world-model imagination. Pre-VLA leverages an efficient multimodal backbone with modality-aware pooling and a lightweight dual-branch head to predict both safety confidence and critic-derived advantage scores for candidate action chunks. To handle severe class imbalance and unstable boundary decisions, we train Pre-VLA with a multi-task objective combining Focal classification, advantage regression, and soft-threshold calibration. During deployment, a dual-mode preemptive resampling scheduler filters low-quality actions and triggers adaptive resampling under a limited computation budget. Experiments on the LIBERO benchmark show that Pre-VLA improves the average closed-loop success rate across four suites from 30.79\% to 37.62\% over RynnVLA-002, reduces task execution steps, achieves 183.9 ms average forward verification time per action chunk, and mitigates error accumulation in world-model rollouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。