arXiv:2602.15882cs.ROcs.AI2026-02被引 3

让机器人实时预测未来轨迹,延迟不变却能看更远。

FUTURE-VLA: Forecasting Unified Trajectories Under Real-time Execution

  • 把长时控制与未来预测统一为序列生成任务,用自适应压缩提升信息密度。
  • 在保持单帧延迟下,实现16倍时空窗口扩展,真实场景成功率超78%。
  • 支持人机交互预览,可动态验证行为,适合具身智能与机器人应用。

通用视觉语言模型日益支持对长视频流的统一时空推理,但将其部署于机器人时受限于处理长时历史和生成高维未来预测的高昂延迟。为此,我们提出FUTURE-VLA,一种统一架构,将长时控制与未来预测重构为单一序列生成任务。采用双侧高效范式,FUTURE-VLA利用时间自适应压缩策略最大化时空信息密度,可在保持恒定推理延迟的前提下处理大规模多视角历史。同时,通过潜在空间自回归,将可执行动力学与可回溯视觉前瞻对齐于一次前向传播。这些实时预测能力进一步支持基于预测的人机协同机制,通过交互式执行门控实现操作员对行为的动态验证。大量评估表明,FUTURE-VLA达到新最优性能,在LIBERO上成功率达99.2%,RoboTwin上达75.4%,真实世界Piper平台达78.0%,且在16倍扩展的时空窗口下仍保持单帧基准推理延迟。

原文摘要 · Abstract (English)

General vision-language models increasingly support unified spatiotemporal reasoning over long video streams, yet deploying such capabilities on robots remains constrained by the prohibitive latency of processing long-horizon histories and generating high-dimensional future predictions. To bridge this gap, we present FUTURE-VLA, a unified architecture that reformulates long-horizon control and future forecasting as a monolithic sequence-generation task. Adopting a dual-sided efficiency paradigm, FUTURE-VLA leverages a temporally adaptive compression strategy to maximize spatiotemporal information density, enabling the ingestion of extensive multi-view histories while maintaining constant inference latency. Simultaneously, it performs latent-space autoregression to align actionable dynamics with reviewable visual look-aheads in a single forward pass. These real-time predictive capabilities further enable a prediction-guided Human-In-the-Loop mechanism via interactive execution gating, allowing operators to dynamically validate behaviors based on interpretable future previews. Extensive evaluations demonstrate that FUTURE-VLA establishes new state-of-the-art performance, attaining success rates of 99.2% on LIBERO, 75.4% on RoboTwin, and 78.0% on a real-world Piper platform, all with a $16\times$ extended spatiotemporal window while maintaining the inference latency of a single-frame baseline.

机器人轨迹预测实时推理视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。