用矩阵乘积态建模时间序列,同时完成分类与补全任务。
Using matrix-product states for time-series machine learning
- 基于矩阵乘积态构造时序数据联合概率模型,直接从数据学习相关性结构。
- 在医学、能源、天文数据上表现媲美顶尖方法,最大纠缠度χ_max=20-160。
- 可同时用于分类与缺失值补全,且输出完整概率分布,适合需要解释性的场景。
矩阵乘积态(MPS)在量子多体物理中已被证明是描述复杂关联的高效近似方法,尤其在一维系统中既能捕捉关键量子关联又可在经典计算机上高效处理。这促使研究者将其引入机器学习领域,以应对数据中复杂相关性的建模需求。本文提出一种基于MPS的算法MPSTime,用于学习观测时间序列数据背后的联合概率分布,并应用于分类与缺失值补全等核心问题。该方法能高效从数据中学习复杂的时序概率结构,仅需适中的最大纠缠度χ_max(20-160),并通过单一对数损失函数实现分类与补全任务的统一训练。我们在合成数据及医疗、能源、天文等真实世界数据集上验证了其性能,结果与当前最优方法相当,且优势在于完整编码了数据的联合概率分布,便于结构分析与解释。论文配套开源代码包MPSTime已发布,展示了该方法在科学、工业和医学中解决复杂时序分析问题的强大潜力。
原文摘要 · Abstract (English)
Matrix-product states (MPS) have proven to be a versatile ansatz for modeling quantum many-body physics. For many applications, and particularly in one-dimension, they capture relevant quantum correlations in many-body wavefunctions while remaining tractable to store and manipulate on a classical computer. This has motivated researchers to also apply the MPS ansatz to machine learning (ML) problems where capturing complex correlations in datasets is also a key requirement. Here, we develop and apply an MPS-based algorithm, MPSTime, for learning a joint probability distribution underlying an observed time-series dataset, and show how it can be used to tackle important time-series ML problems, including classification and imputation. MPSTime can efficiently learn complicated time-series probability distributions directly from data, requires only moderate maximum MPS bond dimension $χ_{\rm max}$, with values for our applications ranging between $χ_{\rm max} = 20-160$, and can be trained for both classification and imputation tasks under a single logarithmic loss function. Using synthetic and publicly available real-world datasets, spanning applications in medicine, energy, and astronomy, we demonstrate performance competitive with state-of-the-art ML approaches, but with the key advantage of encoding the full joint probability distribution learned from the data, which is useful for analyzing and interpreting its underlying structure. This manuscript is supplemented with the release of a publicly available code package MPSTime that implements our approach. The effectiveness of the MPS-based ansatz for capturing complex correlation structures in time-series data makes it a powerful foundation for tackling challenging time-series analysis problems across science, industry, and medicine.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。