通过自验证机制提升机器人动作模型的决策可靠性
World-Coherent Decoding: Self-Verifying Test-Time Planning for World Action Models

- 用生成信号评估未来画面合理性和动作稳定性,动态筛选最优路径
- 实测在有限监督下硬成功率从55.80%提升至60.90%,Horizon-3任务增益16.43%
- 无需更新主模型,适合真实场景中对鲁棒性要求高的机器人部署
世界动作模型(WAMs)通过随机生成视觉未来并解码动作来控制机器人,但结果高度依赖所选未来。本文提出世界一致性解码(WCD),一种测试时规划的自验证框架,将WAM展开视为可检验的未来-动作假设。每个决策步骤中,WCD从冻结的WAM采样多个候选,并利用基于流的视频意外度评估视觉合理性,用动作路径耗散度评估动作生成稳定性进行排序。执行后,实际观测结果验证所选想象,产生想象与现实的偏差,用于训练轻量级在线预测器以优化未来选择。该方法将延迟自验证转化为预执行可靠性估计,无需更新主模型。在RoboTwin 2.0上,WCD使受限随机场景监督下的硬成功率从55.80%提升至60.90%,在Horizon-3任务上取得+16.43%的提升,并在真实Franka视觉偏移测试中展现定性鲁棒性。结果表明:WAM的测试时扩展更关键的是选择可靠未来,而非增加采样数量。
原文摘要 · Abstract (English)
World Action Models (WAMs) aim to control robots by stochastically generating visual futures and then decoding actions, but empirical observations indicate that the results can strongly depend on which future is selected. We propose World-Coherent-Decoding (WCD), a self-verifying test-time planning framework that treats WAM rollouts as falsifiable future--action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using internal generative signals: flow-based video surprisal for visual plausibility and action path effort for action-generation stability. After execution, the realized observation audits the selected imagination, yielding an imagination--reality mismatch that trains a lightweight online predictor for future candidate selection. Thus, WCD converts delayed self-verification into pre-execution reliability estimation without updating the backbone model. On RoboTwin 2.0, WCD improves Hard success under limited randomized-scene supervision from $55.80\%$ to $60.90\%$, with a $+16.43$ gains on Horizon-3 tasks, and shows qualitative robustness on real Franka visual-shift tests. These results highlight a simple principle: test-time scaling for WAMs depends less on sampling more futures than on selecting reliable ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。