arXiv:2606.09630cs.ROcs.AI2026-06被引 1

用视觉语言模型指导恢复策略,让机器人失败后能自动纠错

ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies

论文配图:ReCoVLA: VLM-Guided Reward Compilation for Failure Recovery in Vision-Language-Action Policies
图 1 · 摘自论文原文
  • 用外部VLM识别失败原因和恢复阶段,生成结构化奖励信号
  • 模拟中成功率从36.7%提升至66.7%,物理部署达61.7%成功
  • 无需微调主策略,适配多种视觉语言动作模型

视觉-语言-动作(VLA)策略在语言控制的操控任务中具有强大先验,但在异常状态下方差大,需针对性恢复。本文提出ReCoVLA——一种故障条件下的残差恢复框架,保持预训练VLA策略不变,利用外部视觉语言模型(VLM)推断故障模式与恢复阶段,并从任务相关组件中编译出结构化奖励。ReCoVLA不直接用VLM生成动作或奖励,而是将其作为语义奖励选择器:预测恢复描述符与奖励掩码,用于仿真中的残差策略训练,随后零样本实现仿真到现实的部署。该方法将高层故障理解与低层修正控制解耦,支持不同VLA。在短时、长时及接触密集型操控任务中实验表明,ReCoVLA平均优于基线。模拟中,奖励编译器使成功率从36.7%(fine-tuned $π_{0.5}$)提升至66.7%;物理零样本测试中,达到61.7%平均成功率。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies provide strong priors for language-conditioned manipulation, but remain brittle in off-nominal states requiring targeted recovery. We propose ReCoVLA -- a failure-conditioned residual recovery framework that keeps a pretrained VLA policy frozen, uses an external vision-language model (VLM) to infer the failure mode and recovery stage, and compiles a structured reward from task-relevant components. Rather than using the VLM to generate actions or rewards directly, ReCoVLA uses it as a semantic reward selector: it predicts a recovery descriptor and reward mask for in-simulation residual-policy training, followed by zero-shot sim-to-real deployment of the trained recovery policies. This decouples high-level failure understanding from low-level corrective control to support different VLAs. Experiments across short-horizon, long-horizon, and contact-rich manipulation tasks show that ReCoVLA outperforms the tested baselines on average. In simulation, our reward compiler improves average success from 36.7% for the fine-tuned $π_{0.5}$ baseline to 66.7%. In physical zero-shot sim-to-real experiments, ReCoVLA achieves the best average performance, with 61.7% success.

机器人控制视觉语言模型故障恢复零样本部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。