arXiv:2606.10040cs.RO2026-06被引 9

10亿参数模型实现快速未来预测,让机器人实时决策更高效

Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination

论文配图:Efficient-WAM: A 1B-Parameter World-Action Model with Low-Cost Future Imagination
图 1 · 摘自论文原文
  • 用轻量视频专家和稀疏隐变量降低推理开销
  • 未来预测粗糙但能有效指导动作,延迟降至100毫秒
  • 适合对实时性要求高的真实机器人控制场景

世界-动作模型(WAMs)通过结合未来视觉预测与动作生成,为具身控制提供了新范式。然而现有模型依赖高保真视觉预测,导致推理延迟高,难以实现实时部署。为此,我们提出Efficient-WAM,一种10亿参数的世界-动作模型,在保持控制性能的同时显著降低未来想象成本。该模型通过从WAN-2.2-5B迁移的紧凑视频专家、稀疏视频隐变量以及不对称的视频-动作去噪策略(视频采样步数少于动作)提升效率。不同于追求视觉保真度,Efficient-WAM将未来视频预测视为动作生成的紧凑引导信号。在RoboTwin 2.0及真实操控任务上的实验表明,尽管未来预测外观粗糙,其动作性能依然强劲。在实际部署中,本模型每块处理延迟约100毫秒,相比现有WAMs提速30倍,同时保持了竞争性的控制能力。

原文摘要 · Abstract (English)

World-Action Models (WAMs) have emerged as a promising paradigm for embodied control by coupling future visual prediction with action generation. However, most existing WAMs rely on photorealistic future prediction, which incurs high inference latency and makes real-time robot deployment difficult. This motivates a more efficient WAM design that preserves the control benefits of future visual prediction while reducing its inference cost. We introduce Efficient-WAM, a World-Action Model that reduces the cost of future imagination while preserving its control benefit. Efficient-WAM improves inference efficiency via a compact video expert transferred from WAN-2.2-5B, token-sparse video latents, and asymmetric video-action denoising that allocates fewer sampling steps to video than to actions. Instead of optimizing the future branch for visual fidelity, Efficient-WAM treats future video prediction as a compact guidance signal for action generation. Comprehensive experiments on RoboTwin 2.0 and real-world manipulation tasks show that Efficient-WAM maintains strong action performance despite visibly coarse future predictions. While maintaining competitive control capabilities, our 1B-parameter model can reduce per-chunk latency to around 100 ms during physical deployment, achieving a 30x speedup over existing WAMs.

机器人控制视觉预测高效模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。