提出并行变分蒙特卡洛方法,让深度状态空间模型训练快10倍。
Efficient Learning of Deep State Space Models via Importance Smoothing
- 用并行变分蒙特卡洛融合生成与判别训练范式
- 在基准测试中性能达或超现有最优,训练速度提升10倍
- 适合需要高效训练深度时序模型的研究者
潜在状态空间系统在统计建模中无处不在,尤其在通过噪声观测获取时间序列数据时。然而,大规模训练深度状态空间模型(DSSMs)仍具挑战。目前主要有两种策略:一是通过优化变分下界训练生成模型的自编码类DSSMs;二是对经典顺序蒙特卡洛(SMC)算法输出进行反向传播。这些方法虽可支持判别与生成任务,但其固有的序列前向过程在现代硬件上扩展性差。本文提出并行变分蒙特卡洛(PVMC),一种新训练方法,融合两类范式,可稳健训练用于判别与生成任务的DSSMs。在一系列基准实验中,PVMC性能达到或超过当前最优水平,且训练速度比最快的现有SMC方法快10倍。
原文摘要 · Abstract (English)
Latent state space systems are ubiquitous in statistical modelling, arising naturally when time series are observed through noisy measurements. However, training deep state space models (DSSMs) at scale remains difficult. Two largely distinct strategies have emerged for training DSSMs. The first, auto-encoding DSSMs, trains generative models by optimising a variational lower bound. The second backpropagates through the outputs of classical sequential Monte Carlo (SMC) algorithms. Such approaches can train DSSMs for both discriminative and generative tasks, but their inherently sequential forward passes scale poorly on modern hardware. We propose \emph{parallel variational Monte Carlo} (PVMC), a new training method that bridges these paradigms and robustly trains DSSMs for both discriminative and generative tasks. Across a set of benchmark experiments, PVMC matches or exceeds state-of-the-art performance while training $10\times$ faster than the fastest competing SMC-based approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。