arXiv:2606.01027cs.RO2026-06被引 12

一个能预测未来并评估动作的机器人操作模型,提升长程精细任务表现。

$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation

论文配图:$τ_0$-WM: A Unified Video-Action World Model for Robotic Manipulation
图 1 · 摘自论文原文
  • 统一框架融合动作生成、视频预测与动作评估,共享视频扩散主干网络。
  • 在2.7万小时真实机器人数据上训练,推理时通过重去噪一致性排序动作候选。
  • 适合需要高精度长程操作的机器人场景,如复杂抓取与装配任务。

机器人操作需要能在执行前生成可执行动作并预测其未来后果的模型。我们提出$τ_0$-世界模型($τ_0$-WM),一种集成策略学习、视频预测与动作评估的统一未来预测框架。基于共享的视频扩散主干,$τ_0$-WM提供两种互补接口:第一,视频动作模型从多视角观测、语言指令和机器人状态联合预测未来视觉隐变量与连续动作块;第二,动作条件视频模拟器将候选动作块展开为多视角未来,并预测密集的任务进展分数。模型在约27,300小时真实机器人遥操作数据、UMI式交互数据、第一人称人类视频及回放或失败轨迹上,使用模态特定监督掩码进行训练。推理时,$τ_0$-WM利用测试时计算采样动作候选,通过重去噪一致性排序,并对低质量候选调用模拟器修正。在具有挑战性的长时序与细粒度机器人操作任务中,$τ_0$-WM性能优于其他相关基线。

原文摘要 · Abstract (English)

Robotic manipulation requires models that generate executable actions while anticipating and evaluating their future consequences before physical execution. We present $τ_0$-World Model ($τ_0$-WM), a unified video-action world model that integrates policy learning, video prediction, and action evaluation within a single future-predictive framework. Built on a shared video diffusion backbone, $τ_0$-WM provides two complementary interfaces. First, a video action model jointly predicts future visual latents and continuous action chunks from multi-view observations, language instructions, and robot state. Second, an action-conditioned video simulator rolls out candidate action chunks into multi-view futures and predicts dense task-progress scores. The model is trained on approximately $27{,}300$ hours of real-robot teleoperation, UMI-style interaction, egocentric human videos, and rollout or failure trajectories using modality-specific supervision masks. At inference time, $τ_0$-WM uses test-time computation to sample action candidates, rank them with re-denoising consistency, and invoke simulator-based rectification for low-quality candidates. On challenging long-horizon and fine-grained robotic manipulation tasks, $τ_0$-WM shows superior performance over other relevant baselines.

机器人操作视频生成动作预测扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。