arXiv:2505.23569cs.LGstat.ML2025-05被引 2

无需重建的时序建模新方法,直接学习隐状态动态。

Maximum Likelihood Learning of Latent Dynamics Without Reconstruction

  • 用最大似然训练隐状态模型,避免显式观测映射
  • 在视频数据上精准捕捉非线性随机动态,抗背景干扰
  • 适合需要高质量隐变量表示的下游任务

我们提出一种新的无监督时序建模方法:识别参数化的高斯状态空间模型(RP-GSSM)。该模型通过最大似然学习马尔可夫高斯隐状态,解释不同时刻观测间的统计依赖关系。与对比方法不同,它是合法的概率生成模型;与生成方法不同,它无需从隐状态到观测的显式映射网络,使模型容量更集中于隐状态推断。模型具备精确推断能力(因隐先验联合高斯),同时通过任意非线性神经网络连接观测与隐状态保持表达力。该方法无需人工正则化、辅助损失或优化器调度即可学习任务相关的隐状态。实验表明,在含或不含背景干扰的视频中学习非线性随机动力学方面优于现有方法。结果表明RP-GSSM可作为多种下游应用的基础模型。

原文摘要 · Abstract (English)

We introduce a novel unsupervised learning method for time series data with latent dynamical structure: the recognition-parametrized Gaussian state space model (RP-GSSM). The RP-GSSM is a probabilistic model that learns Markovian Gaussian latents explaining statistical dependence between observations at different time steps, combining the intuition of contrastive methods with the flexible tools of probabilistic generative models. Unlike contrastive approaches, the RP-GSSM is a valid probabilistic model learned via maximum likelihood. Unlike generative approaches, the RP-GSSM has no need for an explicit network mapping from latents to observations, allowing it to focus model capacity on inference of latents. The model is both tractable and expressive: it admits exact inference thanks to its jointly Gaussian latent prior, while maintaining expressivity with an arbitrarily nonlinear neural network link between observations and latents. These qualities allow the RP-GSSM to learn task-relevant latents without ad-hoc regularization, auxiliary losses, or optimizer scheduling. We show how this approach outperforms alternatives on problems that include learning nonlinear stochastic dynamics from video, with or without background distractors. Our results position the RP-GSSM as a useful foundation model for a variety of downstream applications.

时序建模隐状态生成模型最大似然

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。