arXiv:2608.13621cs.AI2026-08

发现时间序列JEPA其实是个隐马尔可夫模型,有明确的状态演化逻辑。

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

论文配图:Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning
图 1 · 摘自论文原文
  • 用状态空间视角解析JEPA的三重角色:推断、转移、观测发射
  • 提出MCJEPA模型,保证多步预测的马尔可夫一致性
  • 适合对自监督学习机制背后的数学结构感兴趣的读者

隐马尔可夫模型(HMM)具备三个核心功能:从观测推断隐藏状态信念、通过马尔可夫转移传播状态、将状态映射回观测空间。我们证明,完整的时间索引预测信息瓶颈联合嵌入预测模型(PIB-VJEPA)展现出相同的计算结构:随机上下文编码器扮演变分滤波分布角色,概率预测器定义潜在状态动态,解码器或逆目标编码器提供发射方向。我们区分了四种逐步强化的对应关系,并给出精确序列级HMM等价性的充分条件。为具体化这一联系,我们引入马尔可夫链JEPA(MCJEPA),用学习到的转移矩阵替代潜在预测器;在有限时齐情形下,矩阵幂确保多步预测满足精确的Chapman-Kolmogorov一致性。离散状态转移、连续状态马尔可夫核及连续时间动力学扩展该构造,而确定性时间JEPA则作为狄拉克核的退化特例出现。我们进一步将预测信息瓶颈学习解释为寻找紧凑的预测状态:压缩促进最小性,残差可预测性检验充分性。受控实验验证了转移组合、滤波解释、已知合成过程中的预测马尔可夫化,以及JEPA潜变量预测与HMM式序列学习的差异。这些结果共同为时间序列JEPA提供了严谨的状态空间解释。

原文摘要 · Abstract (English)

A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagation through a Markov transition, and emission back to observation space. We show that full, time-indexed Predictive Information Bottleneck VJEPA (PIB-VJEPA) exposes the same computational structure: a stochastic context encoder plays the role of an amortized filtering distribution, a probabilistic predictor defines latent-state dynamics, and a decoder, inverse target encoder, or induced implicit conditional supplies the emission direction. We distinguish 4 progressively stronger levels of correspondence and give sufficient conditions for exact sequence-level HMM equivalence. To make the connection concrete, we introduce Markov-Chain JEPA (MCJEPA), which replaces the latent predictor by a learned transition matrix; in the finite time-homogeneous case, matrix powers guarantee exact multi-horizon Chapman--Kolmogorov consistency. Conditioned discrete-state transitions, continuous-state Markov kernels, and continuous-time dynamics extend this construction, while deterministic temporal JEPA appears as a degenerate Dirac-kernel special case. We further interpret predictive information-bottleneck learning as seeking a compact predictive state: compression promotes minimality, while residual predictability tests sufficiency. Controlled experiments support transition composition, the filtering interpretation, predictive Markovization in a known synthetic process, and the distinction between JEPA latent prediction and HMM-style sequence learning. Together, these results give temporal JEPA a principled state-space interpretation.

自监督学习状态空间模型隐马尔可夫表征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。