通过预想的熟悉未来状态,让机械臂在出错后快速恢复动作。
Back to the Familiar Future: Failure Recovery for VLA Policies via Pre-Imagined Milestone Selection

- 提前生成一系列熟悉的未来视觉状态作为恢复目标
- 在出错时选择最合适的预设状态,使系统稳定回归
- 无需微调底层模型,适合工业场景中的可靠控制
视觉-语言-动作(VLA)策略在操作任务中可能偏离预定轨迹,即使任务仍物理可行。此类偏差会将策略带入陌生状态空间,导致直接重规划常引发动作序列失稳。本文提出回溯熟悉未来(B2FF)恢复框架,利用未来视觉条件作为恢复接口。执行前,VLA基于初始观测生成一组熟悉的未来状态(里程碑库)。出错时,一个具备可恢复性判断能力的选择器从该库中选出合适里程碑,并将其固定为视觉目标。这使VLA能稳健地将偏离状态映射回熟悉未来。在注入失败的LIBERO数据集上,于与故障注入时间对齐的恢复时机下,基准VLA成功率由56.3%提升至74.0%,证明预想的里程碑可在不微调低层动作生成器的前提下有效引导恢复。
原文摘要 · Abstract (English)
Vision-language-action (VLA) policies can deviate from nominal trajectories during manipulation, even when tasks remain physically feasible. Recovering from these deviations is challenging, as they push the policy into unfamiliar state spaces where direct re-planning frequently destabilizes action sequences. We propose Back to the Familiar Future (B2FF), a recovery framework for foresight-driven VLAs that leverages future visual conditioning as a recovery interface. Before execution, the VLA generates a milestone bank of familiar future states conditioned on the clean initial observation. At recovery time, a recoverability-aware selector selects a recovery milestone from this bank and enforces it as a fixed visual goal. This enables the VLA to robustly map off-trajectory observations back to a familiar future. On failure-injected LIBERO, under controlled recovery timing aligned with the injected failure, B2FF increases the average success rate of a baseline VLA from 56.3% to 74.0%, demonstrating that pre-imagined milestones can guide recovery without fine-tuning the low-level action generator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。