arXiv:2506.01392cs.ROcs.AI2025-06中稿 · ICLR被引 7

通过稀疏想象提升视觉世界模型规划效率,让机器人实时决策更可行。

Sparse Imagination for Efficient Visual World Model Planning

  • 用随机分组注意力稀疏化视觉世界模型,动态减少处理的标记数。
  • 在保持控制精度前提下,推理速度显著提升,任务表现几乎不变。
  • 适合资源受限的机器人实时规划,尤其适用于最新视觉语言动作模型。

基于世界模型的规划显著提升了复杂环境中的决策能力,使智能体能够模拟未来状态并做出明智选择。然而,这一方法在机器人领域面临严重计算负担,因资源受限而难以应用。为此,我们提出一种稀疏想象(Sparse Imagination)的高效视觉世界模型规划方法,通过减少前向预测中处理的标记数量来提升计算效率。该方法基于使用随机分组注意力策略的稀疏训练视觉世界模型,使模型可根据计算资源灵活调整处理的标记数量。通过在潜在空间回溯中实现稀疏想象,本方法大幅加速了规划过程,同时保持高控制保真度。实验表明,稀疏想象在维持任务性能的同时显著提升推理效率。这一通用视觉规划技术可应用于从简单测试时轨迹优化到复杂现实任务的多种场景,使世界模型在实时环境中部署成为可能。

原文摘要 · Abstract (English)

World model based planning has significantly improved decision-making in complex environments by enabling agents to simulate future states and make informed choices. This computational burden is particularly restrictive in robotics, where resources are severely constrained. To address this limitation, we propose a Sparse Imagination for Efficient Visual World Model Planning, which enhances computational efficiency by reducing the number of tokens processed during forward prediction. Our method leverages a sparsely trained vision-based world model based on transformers with randomized grouped attention strategy, allowing the model to flexibly adjust the number of tokens processed based on the computational resource. By enabling sparse imagination during latent rollout, our approach significantly accelerates planning while maintaining high control fidelity. Experimental results demonstrate that sparse imagination preserves task performance while dramatically improving inference efficiency. This general technique for visual planning is applicable from simple test-time trajectory optimization to complex real-world tasks with the latest VLAs, enabling the deployment of world models in real-time scenarios.

视觉规划世界模型稀疏推理机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。