让视觉语言动作模型在执行中自动发现并纠正错误,无需重训练。
VLA-Corrector: Stage-Aware Observable State Understanding for Prompt-Based Closed-Loop Recovery of Vision-Language-Action Policies

- 基于多视角视觉和动作历史,实时判断任务阶段与风险。
- 在多个干扰下使任务成功率提升显著,接近有特权信息时的表现。
- 适合已有模型想增强鲁棒性的研究者或工程师使用。
长时程机器人操作中,视觉-语言-动作(VLA)策略易受执行偏差影响,因最终任务成功无法提供足够信息诊断失败原因。本文提出一种阶段感知的故障验证与提示恢复框架,可在不更新参数或依赖特权仿真状态的前提下,实现闭环纠错。该框架引入基于可观测历史的可学习验证器,通过时空建模多视角视觉、本体感知状态与已执行动作,联合估计操作进展与执行风险。为提升可解释性,将操作过程划分为接近、对齐、抓取、运输和放置等语义阶段,并识别各阶段特异性失败模式。一旦检测到异常,保留原始指令并生成阶段条件式恢复提示,使同一冻结的VLA策略输出修正动作。在LIBERO与LIBERO Plus上的多轮评估表明,该方法在多种扰动下显著提升闭环可靠性。即使无物体或目标坐标等特权信息,所提验证器的恢复性能也接近特权规则基验证器。结果表明,仅凭可观测的视觉-本体-动作历史即可推断潜在任务状态,为现有VLA策略提供实用故障恢复能力。
原文摘要 · Abstract (English)
Long-horizon robot manipulation with Vision-Language-Action (VLA) policies remains vulnerable to execution-time deviations, as final task success provides little information for diagnosing and correcting failures caused by action noise, object displacement, or goal misalignment. We introduce a stage-aware failure verification and Prompt Recovery framework that enables closed-loop correction of a fixed VLA policy without parameter updates or privileged simulator states. The framework introduces an observable-history-based Learned Verifier that jointly estimates manipulation progress and execution risk by temporally modeling multi-view visual observations, proprioceptive states, and executed actions. To provide interpretable task understanding, we represent manipulation execution through semantic progress stages, including approach, alignment, grasp, transport, and placement, and identify stage-specific failure patterns. Upon detecting abnormal execution, the framework preserves the original instruction and generates a stage-conditioned recovery prompt, allowing the same frozen VLA policy to produce corrective actions. Extensive multi-round evaluations on LIBERO and LIBERO Plus demonstrate that the proposed approach substantially improves closed-loop reliability under diverse perturbations. Without access to privileged object or goal coordinates, the Learned Verifier achieves recovery performance close to that of the privileged rule-based verifier in the evaluated settings. These results show that observable visual-proprioceptive-action history is sufficient to infer latent task states and enable practical failure recovery for existing VLA policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。