用数据驱动方法构建马尔可夫模型状态,提升预测准确性
A Data-Driven Approach to State Construction in Markov Models

- 结合特征选择与无监督学习自动发现状态
- 谱聚类和自组织映射在捕捉数据结构上表现最佳
- 适合需要精准状态划分的建模场景
马尔可夫链是广泛使用的随机过程,用于建模随时间演化的随机事件。模型依赖于从全数据集中提取的状态子集,这些状态需满足转移概率同质性。然而,状态构建常被忽视或依赖先验假设,可能违背同质性要求,降低模型有效性与预测能力。为此,本文将有监督特征选择与无监督学习结合,提出数据驱动的状态构造方法。考察了密度聚类、谱聚类和柯赫嫩自组织映射在无先验条件下识别潜在分组的能力。研究贡献有二:其一,提出融合合适无监督学习技术的方法框架,并设计分类性能与马尔可夫模型精度的评估指标;其二,在实际应用中测试该框架,对比分析显示谱聚类和自组织映射最能捕捉内在结构。结果为实际马尔可夫建模中的状态定义提供了理论与方法指导。
原文摘要 · Abstract (English)
A Markov chain is a widely used stochastic process modelling random events over time. These models are built on subsets of the entire dataset, referred to as states, which are considered to be homogeneous regarding transition probabilities. However, the creation of these states is often disregarded or based on prior assumption, potentially violating the homogeneity requirement and thus decreasing the validity and predictive power of the model. In order to fill this gap, this paper combines supervised feature selection with unsupervised learning techniques for data-driven state construction. Density-based clustering, spectral clustering, and Kohonen self-organizing maps are examined for their ability to identify latent groups without prior assumptions. The contribution of this study is twofold. First, the paper presents a methodological framework for state construction incorporating suitable unsupervised learning techniques, with appropriate measures both for classification performance and Markov model accuracy. Secondly, the framework is tested on an application, resulting in a comparative analysis showing that spectral clustering and Kohonen self-organizing maps are best at capturing inherent structure. These results serve as a cornerstone in providing theoretical and methodological guidance for improving state definition in applied Markov modelling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。