无需训练即可在执行时修复机器人任务中断,提升成功率近85%。
Imagining Recovery: Inference-Time Counterfactual Realignment for Vision-Language-Action Models

- 通过想象理想路径并最小调整状态实现无试错恢复
- 实测成功率达近正常水平,物理重置次数减少42.2%
- 适用于真实场景中指令变更与环境扰动,无需额外数据
视觉-语言-动作(VLA)模型提升了机器人操作的灵活性与通用性,但对在线干扰(如目标变化、场景配置改变或机器人状态异常)仍显脆弱。现有恢复方法常需失败数据、策略重训或外部校正代理,带来额外数据需求和执行风险。本文提出反事实重对齐(CoRe),一种无需训练的推理时恢复框架,在不依赖失败数据的前提下,于推理阶段恢复冻结的VLA模型。当检测到偏离时,CoRe基于近期可行状态,合成观测以想象策略如何继续完成当前目标,并最小化地调整机器人与环境状态,使系统重新接续该预想路径后交还控制权。恢复过程无需物理试错,保留已完成任务进度,统一处理任务中指令变更与物理扰动。跨多个仿真器、不同VLA主干网络及真实场景的大量实验表明,CoRe将成功率提升最高达85.0个百分点至接近正常水平,同时减少42.2%的物理重置次数,且无需策略微调或特定失败场景的恢复训练。
原文摘要 · Abstract (English)
Vision-language-action (VLA) models have improved the flexibility and generality of robotic manipulation, yet they remain fragile to online disruptions, such as changes in task goal, scene configuration, or robot state. Existing recovery methods often require failure data, policy retraining, or external corrective agents, introducing additional data requirements and execution risks. We propose Counterfactual Realignment (CoRe), a training-free framework that recovers a frozen VLA at inference time without failure data. Upon detecting a deviation, CoRe imagines how the policy would continue toward the current goal from a recent viable state, using synthesized observations in place of physical execution, and then minimally realigns the robot and scene to rejoin this imagined continuation before returning control to the policy. Recovery is therefore planned without physical trial-and-error, preserves completed task progress, and handles both mid-episode instruction changes and physical perturbations in a unified manner. Extensive experiments across multiple simulators, VLA backbones, and real-world settings show that CoRe improves success rates by up to 85.0 percentage points to near-nominal levels while reducing physical restorations by 42.2%, without policy fine-tuning or failure-specific recovery training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。