arXiv:2607.04978cs.LGcs.CV2026-07被引 2

Qantara让同一个模型同时支持多种控制推理方式,提升部署灵活性。

Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control

论文配图:Qantara: Bridge-Flow Training for Multi-Paradigm JEPA Control
图 1 · 摘自论文原文
  • 用布朗桥与动作流匹配联合训练,统一多范式推理路径。
  • 在LeWM上达91.2%成功率,刷新OGBench-Cube和LeWM基准。
  • 同一模型可直接用于规划、行为克隆和逆动力学,适合复杂场景部署。

联合嵌入预测架构(JEPAs)正成为从原始像素学习潜空间世界模型的核心方法,但现有模型在训练时仅固定一种推理范式:要么基于学习的动力学模型进行轨迹优化,要么直接行为克隆。本文提出Qantara,一个端到端的JEPA模型,其联合训练目标在状态轴上使用布朗桥插值连续干净潜在表示,在动作轴上实现噪声到数据的流匹配。同一检查点无需重训练即可支持三种推理范式:潜空间规划、行为克隆动作采样与逆动力学推断,通过视频逆向组合实现:先无动作条件预测下一潜状态,再提取动作。训练聚焦于(动作-时间,状态-时间)噪声方块的边缘区域,此处为推理查询提供高置信预测;若改用内部均匀采样,相同计算量下Push-T任务规划成功率从90.1降至53.3。在LeWM控制套件上,Qantara平均达到91.2%成功率,于OGBench-Cube上超越DINO-WM 7.7个百分点、超越LeWM 19.7个百分点。相同权重下,行为克隆与视频逆向路径在Push-T上达成82–83%、在Cube上达71–73%成功率。该工作使JEPA世界模型从单范式规划器跃升为多范式控制器。

原文摘要 · Abstract (English)

Joint-Embedding Predictive Architectures (JEPAs) underpin a growing family of latent world models for control from raw pixels, but every existing JEPA world model commits at training time to a single inference paradigm: either trajectory optimisation in a learned dynamics model, or direct behaviour cloning. A single checkpoint that serves both would defer this choice to inference, when deployment constraints (rollout cost, observation accessibility) determine which path wins. We present Qantara, an end-to-end JEPA whose joint training objective pairs a Brownian-bridge interpolant between consecutive clean latents on the state axis with noise-to-data flow matching on the action axis. The same checkpoint serves three inference paradigms without retraining: latent planning, behaviour-cloning action sampling, and inverse dynamics, which we query through a video-inverse composition that first predicts the next latent without action conditioning, then extracts the action. Training concentrates mass on the edges of the (action-time, state-time) noise square, where inference queries the predictor: replacing it with uniform interior sampling drops Push-T planning from 90.1 to 53.3 SR at matched compute. On the LeWM control suite, Qantara reaches a 91.2 SR three-train-seed average and sets new SOTA on OGBench-Cube (+7.7 SR over DINO-WM, +19.7 over LeWM). From the same weights, the behaviour-cloning and video-inverse paths reach 82-83 SR on Push-T and 71-73 SR on Cube. These results move JEPA world models from single-paradigm planners to multi-paradigm controllers.

世界模型控制多范式扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。