arXiv:2509.12249cs.LGcs.AI2025-09被引 4

辅助任务让JEPA更好区分视觉状态,避免表征坍塌。

Why and How Auxiliary Tasks Improve JEPA Representations

  • 用辅助回归任务与动态一致性联合训练JEPA
  • 理论证明:双损失为零时,不同状态必对应不同表征
  • 适合研究表征学习与强化学习的开发者参考

联合嵌入预测架构(JEPA)在视觉表征学习和基于模型的强化学习中应用日益广泛,但其行为机制仍不明确。本文针对一种简单实用的JEPA变体进行理论分析,该变体包含一个与潜在动态共同训练的辅助回归头。我们证明了‘无有害表征坍塌定理’:在确定性马尔可夫决策过程(MDP)中,若潜在转移一致性损失与辅助回归损失均趋近于零,则任意非等价观测(即具有不同转移动态或辅助值的状态)必然映射到不同的潜在表征。因此,辅助任务定义了表征必须保留的关键区分。在计数环境中的受控消融实验验证了该理论,结果表明联合训练比分别训练产生更丰富的表征。本工作指明了一条改进JEPA编码器的路径:通过与辅助函数联合训练,使潜在表征同时捕捉正确的等价关系。

原文摘要 · Abstract (English)

Joint-Embedding Predictive Architecture (JEPA) is increasingly used for visual representation learning and as a component in model-based RL, but its behavior remains poorly understood. We provide a theoretical characterization of a simple, practical JEPA variant that has an auxiliary regression head trained jointly with latent dynamics. We prove a No Unhealthy Representation Collapse theorem: in deterministic MDPs, if training drives both the latent-transition consistency loss and the auxiliary regression loss to zero, then any pair of non-equivalent observations, i.e., those that do not have the same transition dynamics or auxiliary value, must map to distinct latent representations. Thus, the auxiliary task anchors which distinctions the representation must preserve. Controlled ablations in a counting environment corroborate the theory and show that training the JEPA model jointly with the auxiliary head generates a richer representation than training them separately. Our work indicates a path to improve JEPA encoders: training them with an auxiliary function that, together with the transition dynamics, encodes the right equivalence relations.

表征学习JEPA强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。