用记忆模型让机器人理解过去、预测未来,实现长程智能控制
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
- 构建三系统架构,融合视觉语言与视频扩散模型,形成可记忆的动态世界模型
- 在真实场景中实现36Hz高效运行,长序列任务成功率显著优于基线模型
- 适合需要长期规划与开放指令理解的通用机器人研究者参考
近期视觉语言模型(VLM)进展使机器人能理解开放式指令并展现出色常识推理能力。然而,现有视觉-语言-动作(VLA)框架多依赖静态表征和有限时序上下文,导致代理仅具备短时程、反应式行为,在动态具身环境中难以泛化。受认知神经科学中情景记忆理论启发,我们首次在VLA中形式化构建情景世界模型,使具身机器人能够积累、回溯并预测序列化经验。作为该理念的具体实现,统一的TriVLA采用三系统架构:通过预训练的VLM(System 2)实现多模态定位,利用视频扩散模型(System 3)捕捉丰富的时序动态。这使得代理可积累并回溯经验,解析当前上下文,并预测环境演化。下游策略(System 1)基于跨越过去与未来的情景表征,通过流匹配与跨模态注意力机制生成连贯、情境感知的动作序列。实验表明,TriVLA以约36 Hz的效率运行,在标准基准与挑战性真实世界操作任务中持续优于基线模型,展现出强大的长时程规划与开放意图理解能力,验证了情景世界模型驱动推理对鲁棒、可泛化机器人智能的优势。项目页面:https://zhenyangliu.github.io/TriVLA/
原文摘要 · Abstract (English)
Recent advances in vision-language models (VLMs) have enabled robots to follow open-ended instructions and demonstrate impressive commonsense reasoning. However, current vision-language-action (VLA) frameworks primarily rely on static representations and limited temporal context, restricting agents to short-horizon, reactive behaviors and hindering robust generalization in dynamic embodied environments. Inspired by cognitive neuroscience theories of episodic memory, we propose, to our knowledge, one of the first formalized episodic world models in VLA, enabling embodied robots to accumulate, recall, and predict sequential experiences. As an instantiation of this concept, our unified TriVLA realizes the episodic world model through a triple-system architecture: integrating multimodal grounding from a pretrained VLM (System 2) and temporally rich dynamics perception from a video diffusion model (System 3). This enables the agent to accumulate and recall sequential experiences, interpret current contexts, and predict future environmental evolution. Guided by episodic representations that span both the past and anticipated future, the downstream policy (System 1) generates coherent, context-aware action sequences through flow-matching and cross-modal attention mechanisms. Experimental results show that TriVLA operates efficiently at approximately 36 Hz and consistently outperforms baseline models on standard benchmarks and challenging real-world manipulation tasks. It demonstrates strong long-horizon planning and open-ended intent understanding, showcasing the advantages of episodic world model-inspired reasoning for robust, generalizable robot intelligence. Project Page: https://zhenyangliu.github.io/TriVLA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。