将世界模型的生成计算内化为当前状态表示,实现超高效机器人控制。
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

- 通过多层级中间状态监督当前编码器,将未来生成过程的计算内化到当前表征中。
- 在真实机器人任务中将动作延迟降低3.7倍以上,最高达10.1倍。
- 能自动适应环境变化,适合对实时性要求高的具身智能应用。
世界生成模型通常通过其输出使用:渲染的未来画面、视频条件动作或由昂贵生成分支计算出的潜在上下文。我们提出,其更具可复用的价值在于生成过程中产生的计算。当生成器将受损的未来转化为连贯轨迹时,其中间状态组织了外观、空间布局与跨抽象层次的交互。能否将这种未来生成计算内化为仅从当前视觉上下文推断的表征?我们提出Enfold,将这一计算转移至从当前视觉与语言指令中预测的表征中。训练时,生成器处理观测未来过程中暴露的多层级状态,监督仅基于当前信息的编码器。学习到的表征被反馈以条件生成未来,并由任务头读取,但任务梯度不回传重塑编码器。部署时,动作预测不再执行生成器。在LIBERO、RoboTwin2.0及真实机器人任务中,Enfold支持强控制性能,同时相比Fast-WAM将动作延迟降低3.7倍,Enfold-Flash更达10.1倍。表征分析显示,该表征抑制了无关扰动,优先捕获长时程变化。当当前场景受人为干预改变时,生成延续与执行动作均自适应调整,这与固定轨迹重放不一致。这些结果重新定义了世界生成器:其未来无需每步生成,只要内部结构可折叠进当前状态即可。
原文摘要 · Abstract (English)
World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。