arXiv:2603.00049cs.LG2026-03被引 2

双向预测提升表征学习,解决传统模型信息损失问题。

BiJEPA: Bi-directional Joint Embedding Predictive Architecture for Symmetric Representation Learning

  • 采用双向预测机制,同时学习正反向数据关系。
  • 在混沌系统和图像数据上实现稳定收敛,避免表示崩溃。
  • 适合需要对称表征的时序与多模态学习任务。

自监督学习已从像素级重建转向潜在空间预测,以联合嵌入预测架构(JEPA)为代表。然而,标准JEPA模型通常依赖单向预测(如上下文→目标),可能忽略逆向关系中的信息,影响性能。本文提出双向联合嵌入预测架构(BiJEPA),强制数据片段间的循环一致性预测。为应对对称预测固有的不稳定性(表示爆炸),引入关键的范数正则化机制。我们在三类不同模态上评估:合成周期信号、洛伦兹吸引子轨迹以及高维图像数据(MNIST)。结果表明,BiJEPA实现了无崩溃的稳定收敛,捕捉了混沌系统的语义结构,并学习到具备生成与泛化能力的鲁棒时空表示,提供更全面的表征学习方法。

原文摘要 · Abstract (English)

Self-Supervised Learning (SSL) has shifted from pixel-level reconstruction to latent space prediction, spearheaded by the Joint Embedding Predictive Architecture (JEPA). While effective, standard JEPA models typically rely on a uni-directional prediction mechanism (e.g. Context $\to$ Target), potentially neglecting the informative signal inherent in the inverse relationship, degrading its performance. In this work, we propose \textbf{BiJEPA}, a \textit{Bi-Directional Joint Embedding Predictive Architecture} that enforces cycle-consistent predictability between data segments. We address the inherent instability of symmetric prediction (representation explosion) by introducing a critical norm regularization mechanism on the representation vectors. We evaluate BiJEPA on three distinct modalities: synthetic periodic signals, chaotic Lorenz attractor trajectories, and high-dimensional image data (MNIST). Our results demonstrate that BiJEPA achieves stable convergence without collapse, captures the semantic structure of chaotic systems, and learns robust temporal and spatial representations capable of generation and generalisation, offering a more holistic approach to representation learning.

自监督学习表征学习双向预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。