arXiv:2604.13518cs.LGcs.AI2026-04被引 2

提出预测性表征学习新范式,突破传统自监督局限

From Alignment to Prediction: A Study of Self-Supervised Learning and Predictive Representation Learning

论文配图:From Alignment to Prediction: A Study of Self-Supervised Learning and Predictive Representation Learning
图 1 · 摘自论文原文
  • 定义预测性表征学习(PRL),以未观测数据为预测目标
  • MAE相似度达1.00但鲁棒性仅0.55,JEPA鲁棒性达0.78
  • 揭示预测建模是未来自监督学习的关键方向

自监督学习已成为从无标签数据中学习的主要技术,现有方法主要围绕表征对齐与输入重构展开。尽管这些方法在实践中表现优异,其应用范围仍局限于可观测数据,难以构建具有预测能力的学习结构。本文研究了自监督学习的最新进展,提出一类新范式——预测性表征学习(PRL),其核心是基于观测数据预测未观测部分。我们构建了一个统一分类体系,将PRL与对齐、重构类方法并列。进一步论证联合嵌入预测架构(JEPA)是该范式的典型代表。通过实现BYOL、MAE和Image-JEPA进行对比分析,结果表明:MAE相似度为1.00,但鲁棒性仅为0.55;BYOL与I-JEPA分别达到0.98与0.95的准确率,鲁棒性分别为0.75与0.78。

原文摘要 · Abstract (English)

Self-supervised learning has emerged as a major technique for the task of learning from unlabeled data, where the current methods mostly revolve around alignment of representations and input recon struction. Although such approaches have demonstrated excellent performance in practice, their scope remains mostly confined to learning from observed data and does not provide much help in terms of a learning structure that is predictive of the data distribution. In this paper, we study some of the recent developments in the realm of self-supervised learning. We define a new category called Predictive Representation Learning (PRL), which revolves around the latent prediction of unobserved components of data based on the observation. We propose a common taxonomy that classifies PRL along with alignment and reconstruction-based learning approaches. Furthermore, we argue that Joint-Embedding Predictive Architecture(JEPA) can be considered as an exemplary member of this new paradigm. We further discuss theoretical perspectives and open challenges, highlighting predictive representation learning as a promising direction for future self-supervised learning research. In this study, we implemented Bootstrap Your Own Latent (BYOL), Masked Autoencoders (MAE), and Image-JEPA (I-JEPA) for comparative analysis. The results indicate that MAE achieves perfect similarity of 1.00, but exhibits relatively weak robustness of 0.55. In contrast, BYOL and I-JEPA attain accuracies of 0.98 and 0.95, with robustness scores of 0.75 and 0.78, respectively.

自监督学习预测表征模型比较

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。