arXiv:2605.25620cs.AI2026-05

从视觉基础模型中提炼简洁任务相关状态,提升规划控制精度

Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations

论文配图:Back to Parsimonious Latents: Learning Task-Centric World Models from Visual Foundations
图 1 · 摘自论文原文
  • 用线性投影压缩视觉嵌入,构建紧凑动态空间
  • 对比学习对齐物理状态子空间,重建保留有用视觉结构
  • 适合无奖励离线强化学习场景,跨环境规划表现更优

世界模型通过动作条件预测未来动态,其潜在表示的选择直接影响规划与控制效果。现有方法或直接从像素学习语义结构有限的表示,或继承冻结视觉基础模型中冗余的无关细节,导致状态空间与下游任务不匹配,尤其在无奖励离线设置下,模型需从固定轨迹中学习而无奖励监督或在线交互。为此,我们提出TC-WM框架,将基础模型嵌入转化为紧凑、任务充分的世界表示。核心设计是将预训练嵌入空间视为语义骨架而非最终状态空间:通过线性投影将高维视觉嵌入压缩为紧凑潜在空间作为动态空间,利用对比学习对齐子空间与智能体物理状态,并重建嵌入以保留有用视觉结构。该方法结合了基础特征的泛化性与任务中心动态的可控性。理论上,我们证明TC-WM能识别底层任务中心潜在因子,至多一个简单变换。实验上,TC-WM在多种环境(如Robomimic和D4RL)中实现测试时规划,世界建模质量与控制精度均优于现有最优方法。

原文摘要 · Abstract (English)

World models enable agents to predict future dynamics conditioned on actions, making the choice of latent representation central to planning and control. Such representations are often either learned directly from pixels with limited semantic structure or inherited from frozen visual foundation models with excessive task-irrelevant detail, yielding state spaces that are poorly matched to downstream planning and control. This is especially challenging in reward-free offline settings, where the model must learn from fixed trajectories without reward supervision or online interaction. To address this, we propose TC-WM, a framework for turning foundation-model embeddings into compact, task-sufficient world representations. The key design is to treat the pretrained embedding space as a semantic scaffold rather than as the final state space: TC-WM linearly projects high-dimensional visual embeddings into a compact latent as the dynamic space, aligns a subspace with the agent's physical state via contrastive learning, and reconstructs embeddings to preserve useful visual structure. This combines the generality of foundation features with the controllability of task-centric dynamics. Theoretically, we show that TC-WM suffices to identify the underlying task-centric latent factors up to a simple transformation. Empirically, TC-WM enables test-time planning across diverse environments (e.g., Robomimic and D4RL), achieving better world-modeling quality and more precise control than state-of-the-art approaches.

世界模型任务相关视觉基础模型离线强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。