提出新型正则化方法,让模型学习更稀疏高效的图像表示。
Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
- 用截断广义高斯分布替代原有高斯分布,实现对稀疏性的显式控制。
- 在图像分类任务上达到与现有方法相当的性能,同时表示更稀疏非负。
- 适合关注模型压缩、高效表示学习的研究者和工程师。
联合嵌入预测架构(JEPA)通过投影分布匹配学习视角不变表示并防止表征坍缩。现有方法将表示正则化为各向同性高斯分布,但天然偏好稠密表示,无法捕捉高效表示中的稀疏特性。本文提出修正分布匹配正则化(RDMReg),一种切片两样本分布匹配损失,使表示对齐于修正广义高斯(RGG)分布。RGG通过截断机制可显式控制期望ℓ₀范数,其连续截断部分在给定ℓₚ范数和支撑集约束下具有最大熵特性。在JEPA中引入RDMReg得到矩形LpJEPA,严格推广了以往基于高斯的JEPA。实验表明,矩形LpJEPA学习到稀疏、非负的表示,在图像分类基准上表现优异,证明RDMReg可在保留任务相关信息的同时有效强制稀疏性。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for collapse prevention. Existing approaches regularize representations towards isotropic Gaussian distributions, but inherently favor dense representations and fail to capture the key property of sparsity observed in efficient representations. We introduce Rectified Distribution Matching Regularization (RDMReg), a sliced two-sample distribution-matching loss that aligns representations to a Rectified Generalized Gaussian (RGG) distribution. RGG enables explicit control over expected $\ell_0$ norm through rectification, while its continuous truncated component admits a maximum-entropy characterization under expected $\ell_p$ norm and support constraints. Equipping JEPAs with RDMReg yields Rectified LpJEPA, which strictly generalizes prior Gaussian-based JEPAs. Empirically, Rectified LpJEPA learns sparse, non-negative representations with favorable sparsity--performance trade-offs and competitive downstream performance on image classification benchmarks, showing that RDMReg can enforce sparsity while preserving task-relevant information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。