文本预测需保留多种合理补全,而非压缩成单一向量。
The JEPA Paradox in Language: The Geometry of Linguistic Alternatives

- 提出语言中预测目标应保持多解性,避免单点压缩
- 实验发现文本预测会引发表示崩溃与性能下降
- 适合研究自回归与非自回归模型的对比设计
联合嵌入预测架构(JEPAs)在图像、视频和音频中表现良好,但确定性JEPA式隐变量预测尚未成为文本编码器的标准目标。本文指出,这源于平方误差预测与语言条件结构之间的不匹配。核心要求是条件集中性:给定上下文和目标位置,目标表示应集中在单一有意义点。局部图像预测因空间连续性常满足此条件,而掩码文本可存在多个有效补全,其表示未必共享一致中心。通过三个条件——可预测性、不坍塌性与低条件方差——形式化该不匹配,并揭示其失败导致中心退化与坍塌压力。匹配的I-JEPA与T-JEPA实验显示,互信息饱和与目标方差升高先于训练-验证不稳定、有效秩退化、余弦坍塌及下游迁移性能下降。这一模式在五个独立数据种子下重复出现,说明非采样偏差。结果不否定语言预测学习,而是表明文本兼容的JEPA目标必须保留多种合理补全,而非压缩为单一隐变量点。
原文摘要 · Abstract (English)
Joint-Embedding Predictive Architectures (JEPAs) are effective for images, video, and audio, yet deterministic JEPA-style latent prediction has not become a standard objective for text encoders. We argue that this gap reflects a mismatch between squared-error latent prediction and the conditional structure of language. The key requirement is conditional concentration: given a context and target location, the target representation should lie near a single meaningful point. Local image prediction often satisfies this through spatial continuity, whereas masked text can admit multiple valid token or span completions whose representations need not share a coherent center. We formalize this mismatch through three conditions---predictability, non-collapse, and low conditional variance---and show how their failure creates centroid degeneracy and collapse pressure in text. Matched I-JEPA and T-JEPA experiments reveal the predicted sequence: mutual-information saturation and elevated target variance precede train--validation instability, effective-rank degeneration, cosine collapse, and poor downstream transfer. The same pattern appears across five independent data seeds, indicating that it is not a sampling artifact. These results do not rule out predictive learning for language; they show that text-compatible JEPA objectives must preserve multiple plausible completions rather than compress them into a single latent point.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。