提出新方法精准计算隐马尔可夫模型的后验统计分布与改进解码性能。
Advanced posterior analyses of hidden Markov models: finite Markov chain imbedding and hybrid decoding
- 用有限马尔可夫链嵌入法计算隐藏状态的访问次数、停留时间等后验分布。
- 混合解码在多个数据集上优于维特比和后验解码,提升解码准确率。
- 提供可复现代码,适合需要精确推断与高精度解码的研究者使用。
隐马尔可夫模型(HMM)应用中的两大核心任务是:(i) 计算隐藏状态序列的汇总统计量后验分布,(ii) 解码隐藏状态序列。本文提出有限马尔可夫链嵌入(FMCI)方法,用于计算如状态访问次数、总停留时间、停留时长及最长连续段长度等统计量的后验分布,通过条件于观测序列的隐藏状态模拟建立框架。第二部分提出混合分割(hybrid segmentation)以改进HMM解码,实证表明其性能优于维特比解码(Viterbi)与后验解码(posterior decoding),并引入一种新型调参策略。此外,基于加权几何平均,给出了混合损失函数的另一起源推导。我们在多个经典数据集上验证了FMCI与混合解码的有效性,并提供配套代码以确保可复现性。
原文摘要 · Abstract (English)
Two major tasks in applications of hidden Markov models are to (i) compute distributions of summary statistics of the hidden state sequence, and (ii) decode the hidden state sequence. We describe finite Markov chain imbedding (FMCI) and hybrid decoding to solve each of these two tasks. In the first part of our paper we use FMCI to compute posterior distributions of summary statistics such as the number of visits to a hidden state, the total time spent in a hidden state, the dwell time in a hidden state, and the longest run length. We use simulations from the hidden state sequence, conditional on the observed sequence, to establish the FMCI framework. In the second part of our paper we apply hybrid segmentation for improved decoding of a HMM. We demonstrate that hybrid decoding shows increased performance compared to Viterbi or Posterior decoding (often also referred to as global or local decoding), and we introduce a novel procedure for choosing the tuning parameter in the hybrid procedure. Furthermore, we provide an alternative derivation of the hybrid loss function based on weighted geometric means. We demonstrate and apply FMCI and hybrid decoding on various classical data sets, and supply accompanying code for reproducibility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。