arXiv:2607.26789cs.RO2026-07被引 1

用动作条件世界模型实时验证长程机械臂操作,提升执行可靠性。

CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation

论文配图:CheckVLA: Execution-Time Verification with Action-Conditioned World Model for Long-Horizon Mobile Manipulation
图 1 · 摘自论文原文
  • 引入动作条件世界模型,在执行中动态验证预期变化是否发生。
  • 在RoboCasa365上成功率提升至36.1%,较周期重规划高8.5个百分点。
  • 适合需要高可靠性的长程移动操作任务,尤其关注执行容错的场景。

视觉-语言-动作(VLA)策略常通过开环动作块执行长程移动操作,发出多个动作后不再接收高层视觉输入。一个已提交的动作块意味着观察应按预期演变,但意外偏差可能违反此预期,而后续动作仍会传播错误:提交时的策略置信度无法响应后续发生的偏差,仅依赖观测的异常评分也缺乏动作条件参考,难以区分预期变化与未解释的变化。本文提出CheckVLA,利用独立训练、冻结的动作条件世界模型进行执行验证。通过符合校准的风险阈值控制整体干预概率,决定何时干预;超出阈值的程度决定重写后缀对被替换块的保留强度;考虑延迟的硬前缀机制限制仅可替换仍可部署的动作;事件驱动的关键帧库则跨修复保留先前进展的证据。在RoboCasa365上,采用相同训练方案与调用预算,CheckVLA平均成功率达36.1%,优于周期重规划的27.6%(+8.5点)。在5%的误报率目标下,动作条件验证的及时召回率达77.9%,远高于仅观测控制的48.6%和动作随机化控制的37.9%。仿真结果支持动作条件验证可在动作块执行中恢复反馈,同时保持修复与推理延迟一致。

原文摘要 · Abstract (English)

Vision-language-action (VLA) policies commonly execute long-horizon mobile manipulation through open-loop action chunks, issuing multiple actions without receiving new high-level visual input. A committed chunk therefore implies how observations should evolve, but accidental deviations can violate this expectation while the remaining actions continue to propagate the error: commit-time policy confidence cannot react to a deviation that occurs after dispatch, and observation-only anomaly scores lack an action-conditioned reference for separating expected effects from unexplained changes. We propose CheckVLA, which verifies execution with a separately trained, frozen action-conditioned world model. A conformally calibrated risk threshold bounds the episode-level probability of an unnecessary first intervention and determines when to intervene, its exceedance controls how strongly the rewritten suffix retains the superseded chunk, latency-aware hard prefixing restricts replacement to actions that remain deployable, and an event-driven keyframe bank preserves evidence of prior progress across repairs. On RoboCasa365, under a common training recipe and a matched invocation budget, CheckVLA attains a 36.1% average success rate against 27.6% for periodic replanning (+8.5 points). At a matched 5% episode-level false-alarm target, action conditioning raises timely recall to 77.9%, against 48.6% for an observation-only control and 37.9% for an action-shuffled control. These simulation results support action-conditioned verification as a way to restore feedback during chunked execution while keeping the repair consistent with inference latency.

机器人长程操作验证机制世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。