用动作条件世界模型实时验证长程机械臂操作,提升执行可靠性。
CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

- 引入动作条件世界模型,在执行中动态验证预期变化是否发生。
- 在RoboCasa365上成功率提升至36.1%,较周期重规划高8.5个百分点。
- 适合需要高可靠性的长程移动操作任务,尤其关注执行容错的场景。
视觉-语言-动作(VLA)策略常通过开环动作块执行长程移动操作,发出多个动作后不再接收高层视觉输入。一个已提交的动作块意味着观察应按预期演变,但意外偏差可能违反此预期,而后续动作仍会传播错误:提交时的策略置信度无法响应后续发生的偏差,仅依赖观测的异常评分也缺乏动作条件参考,难以区分预期变化与未解释的变化。本文提出CheckVLA,利用独立训练、冻结的动作条件世界模型进行执行验证。通过符合校准的风险阈值控制整体干预概率,决定何时干预;超出阈值的程度决定重写后缀对被替换块的保留强度;考虑延迟的硬前缀机制限制仅可替换仍可部署的动作;事件驱动的关键帧库则跨修复保留先前进展的证据。在RoboCasa365上,采用相同训练方案与调用预算,CheckVLA平均成功率达36.1%,优于周期重规划的27.6%(+8.5点)。在5%的误报率目标下,动作条件验证的及时召回率达77.9%,远高于仅观测控制的48.6%和动作随机化控制的37.9%。仿真结果支持动作条件验证可在动作块执行中恢复反馈,同时保持修复与推理延迟一致。
原文摘要 · Abstract (English)
Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。