用张量方法加速因子化隐马尔可夫模型,让大数据分析更快更省资源。
Tensorized algorithms and scalable filtering methods for hidden Markov and factorial hidden Markov models
- 直接利用张量结构处理多因子状态,避免构建巨大状态空间
- 计算效率大幅提升,支持大规模系统与数据集分析
- 适合需要高效处理多因素时序数据的研究者与工程师
时序数据分析常用隐马尔可夫模型(HMM),但许多现实系统受多个独立因素影响,更宜用因子化隐马尔可夫模型(fHMM),由多个隐马尔可夫链共同生成观测数据。尽管fHMM更具表现力,却可重写为等价的HMM,导致状态空间急剧膨胀,显著增加计算成本。尤其是前向滤波算法在评估、解码和估计任务中变得难以承受,即使对小型系统也是如此。本文提出基于张量代数的可扩展方法,直接利用fHMM的多维结构,无需构造中间的HMM表示。新滤波方法大幅提高计算性能,使大规模系统和数据集的高效分析成为可能,拓展了fHMM的应用范围,为数据密集型应用提供了实用框架。
原文摘要 · Abstract (English)
A common method for the representation and analysis of time-series data is the hidden Markov model (HMM), where each observation is associated with a hidden state that evolves over time. However, many real-world systems are influenced by multiple independent factors, which are more naturally represented by factorial hidden Markov models (fHMM), where several hidden Markov chains jointly generate the observed data. Although an fHMM provides a richer and more realistic representation of many real-world systems, it can be reformulated as an equivalent HMM, but with a significantly larger state-space, leading to a severe increase in computational cost. In particular, the forward filtering algorithm, which is central to evaluation, decoding, and estimation tasks, becomes prohibitively expensive even for small systems. This work focuses on developing scalable methods for time-series analysis using tensor algebra to exploit the multidimensional structure of fHMM directly, without constructing intermediate HMM representations. Our novel filtering approach significantly improves computational performance and enables the efficient analysis of large systems and datasets, extending the scope of fHMM and providing a practical framework for data intensive applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。