用视觉语言模型增强潜空间世界模型的长程语义预测能力
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
- 双路径结构:密集帧建模+均匀采样语言模型推理
- 长时序轨迹预测误差降低18.3%,滚动预测更稳定
- 适合需要理解复杂动作语义的任务,如机器人操控
近期潜空间世界模型(如V-JEPA2)在从视频观测中预测未来状态方面展现出潜力。然而,短时窗内的密集预测限制了时间上下文,使预测偏向局部低层外推,难以捕捉长时序语义,降低下游应用价值。视觉-语言模型(VLM)虽能通过均匀采样帧提供强语义基础和通用知识,但因计算驱动的稀疏采样、语言输出瓶颈及数据范式不匹配,不适合作为独立密集预测器。本文提出一种受VLM引导的JEPA式潜空间建模框架,采用双时间路径:密集JEPA分支建模精细运动与交互线索,均匀采样的VLM「思考者」分支以更大时间步长提供知识丰富指导。为有效传递VLM的渐进推理信号,引入层次金字塔表示提取模块,将多层VLM表示聚合为兼容潜空间预测的引导特征。手部操作轨迹预测实验表明,该方法优于强基线VLM模型与JEPA预测器,且在长时序滚动中表现更鲁棒。
原文摘要 · Abstract (English)
Recent progress in latent world models (e.g., V-JEPA2) has shown promising capability in forecasting future world states from video observations. Nevertheless, dense prediction from a short observation window limits temporal context and can bias predictors toward local, low-level extrapolation, making it difficult to capture long-horizon semantics and reducing downstream utility. Vision--language models (VLMs), in contrast, provide strong semantic grounding and general knowledge by reasoning over uniformly sampled frames, but they are not ideal as standalone dense predictors due to compute-driven sparse sampling, a language-output bottleneck that compresses fine-grained interaction states into text-oriented representations, and a data-regime mismatch when adapting to small action-conditioned datasets. We propose a VLM-guided JEPA-style latent world modeling framework that combines dense-frame dynamics modeling with long-horizon semantic guidance via a dual-temporal pathway: a dense JEPA branch for fine-grained motion and interaction cues, and a uniformly sampled VLM \emph{thinker} branch with a larger temporal stride for knowledge-rich guidance. To transfer the VLM's progressive reasoning signals effectively, we introduce a hierarchical pyramid representation extraction module that aggregates multi-layer VLM representations into guidance features compatible with latent prediction. Experiments on hand-manipulation trajectory prediction show that our method outperforms both a strong VLM-only baseline and a JEPA-predictor baseline, and yields more robust long-horizon rollout behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。