arXiv:2604.09331cs.LGcs.SY2026-04

用物理模型提升视频数据中隐状态学习的稳定性与可解释性。

Stability Enhanced Gaussian Process Variational Autoencoders

  • 基于LTI系统定义构造先验,结合概率与物理规律建模隐变量。
  • 通过半收缩系统参数化,实现无约束优化并避免数值发散。
  • 适合需稳定推断动态系统的视频分析任务,尤其关注可解释性。

提出一种新型稳定增强型高斯过程变分自编码器(SEGP-VAE),用于从高维视频数据间接训练低维线性时不变(LTI)系统。该方法基于LTI系统的定义,推导出新的SEGP先验的均值和协方差函数,使模型能够利用融合概率与可解释物理机制的框架捕捉间接观测的潜在过程。通过完全且无约束的参数化方式,将LTI参数搜索空间限制在半收缩系统集合内,从而支持无约束优化算法训练,并有效避免由非Hurwitz状态矩阵引发的数值问题。案例研究应用SEGP-VAE于螺旋粒子视频数据集,验证了该方法的优势及针对特定任务的设计选择对精确潜状态预测的关键作用。

原文摘要 · Abstract (English)

A novel stability-enhanced Gaussian process variational autoencoder (SEGP-VAE) is proposed for indirectly training a low-dimensional linear time invariant (LTI) system, using high-dimensional video data. The mean and covariance function of the novel SEGP prior are derived from the definition of an LTI system, enabling the SEGP to capture the indirectly observed latent process using a combined probabilistic and interpretable physical model. The search space of LTI parameters is restricted to the set of semi-contracting systems via a complete and unconstrained parametrisation. As a result, the SEGP-VAE can be trained using unconstrained optimisation algorithms. Furthermore, this parametrisation prevents numerical issues caused by the presence of a non-Hurwitz state matrix. A case study applies SEGP-VAE to a dataset containing videos of spiralling particles. This highlights the benefits of the approach and the application-specific design choices that enabled accurate latent state predictions.

变分自编码器高斯过程系统辨识视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。