无需解码器的未来特征预测,提升机器人导航效率
DF$^3$: World Modeling via Decoder-Free Feature Forecasting in Autonomous Navigation

- 在冻结视觉模型中注入可学习查询,直接预测未来特征
- 采用轻量级融合机制,实现粗略运动对齐与精细特征匹配
- 无需解码器,支持感知与控制一体化,适合实时系统
从视频序列中预测未来状态是自主机器人系统的关键挑战,也是世界建模的核心目标。以往基于像素的生成方法过度关注无关细节,导致计算开销巨大;而基于潜在空间的方法虽有所改进,仍依赖重型解码器进行状态到任务的映射,成为计算瓶颈。本文提出无解码器特征预测(DF³),完全在潜在空间建模世界演化,并直接生成任务输出。具体而言,DF³将可学习的空间查询注入冻结视觉基础模型的末尾模块,直接提取未来状态表征。通过轻量级统一的运动感知上下文融合(MACF)机制,无缝整合粗粒度光流扭曲与细粒度潜在交叉相关性,使查询与历史标记表示交互,显式对齐并预测下一帧特征。随后,专用的任务查询探测这些预测特征以完成下游任务。在多个公开基准上的实验及机器人模拟器中的零样本部署表明,DF³性能媲美最先进方法,同时具备更高效率和灵活性,适用于感知与控制一体化系统。
原文摘要 · Abstract (English)
Forecasting future states from video sequences is a critical challenge for autonomous robotic systems and a fundamental objective of world modeling. Prior generative methods operating at the pixel level inevitably overemphasize task-irrelevant details, leading to prohibitive computational overhead. While latent-based approaches attempt to mitigate this by predicting features directly, the persistent reliance on heavy decoders for state-to-task mapping remains a computational bottleneck. In this work, we propose Decoder-Free Feature Forecasting (DF$^3$), a novel framework that models world evolution entirely within the latent space and directly derives task outputs, completely eliminating the need for a decoder. Specifically, DF$^3$ injects learnable spatial queries into the terminal blocks of a frozen vision foundation model to extract future state representations directly. By employing a lightweight, unified Motion-Aware Context Fusion (MACF) mechanism that seamlessly integrates coarse flow warping with fine-grained latent cross-correlation, these queries interact with historical token representations to explicitly align and forecast the feature of the next frame. Subsequently, a specialized set of task queries probes these forecasted features for the downstream task. Extensive experiments on public benchmarks and zero-shot deployment in a robotic simulator demonstrate that DF$^3$ achieves performance comparable to state-of-the-art methods while offering superior efficiency and flexibility for integrated perception and control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。