提出可处理重尾数据的时变图学习方法,提升金融等复杂场景建模精度。
Time-Varying Graph Learning for Data with Heavy-Tailed Distribution
- 用带特定谱特性的图结构增强聚类能力
- 结合学生t分布与非负VAR模型捕捉动态变化
- 支持噪声与缺失值,适合金融等重尾数据
图模型能有效刻画网络数据的潜在结构。现实网络拓扑常随时间变化,学习这种动态交互称为时变图学习。现有方法对异常值不鲁棒,难以处理常见于真实数据(如金融数据)的重尾分布。本文提出一种能高效表示重尾数据的时变图学习方法。不同于传统方法,我们引入具有特定谱特性的图结构以增强数据聚类效果。所提方法基于随机模型:通过非负向量自回归(VAR)模型捕捉图的动态变化,用学生t分布建模源自该时变图的信号。设计了半在线框架下的迭代算法,仅用小批量数据更新图结构。在合成与真实数据上的实验表明,该模型在分析重尾数据(特别是金融市场数据)方面具有显著优势。
原文摘要 · Abstract (English)
Graph models provide efficient tools to capture the underlying structure of data defined over networks. Many real-world network topologies are subject to change over time. Learning to model the dynamic interactions between entities in such networks is known as time-varying graph learning. Current methodology for learning such models often lacks robustness to outliers in the data and fails to handle heavy-tailed distributions, a common feature in many real-world datasets (e.g., financial data). This paper addresses the problem of learning time-varying graph models capable of efficiently representing heavy-tailed data. Unlike traditional approaches, we incorporate graph structures with specific spectral properties to enhance data clustering in our model. Our proposed method, which can also deal with noise and missing values in the data, is based on a stochastic approach, where a non-negative vector auto-regressive (VAR) model captures the variations in the graph and a Student-t distribution models the signal originating from this underlying time-varying graph. We propose an iterative method to learn time-varying graph topologies within a semi-online framework where only a mini-batch of data is used to update the graph. Simulations with both synthetic and real datasets demonstrate the efficacy of our model in analyzing heavy-tailed data, particularly those found in financial markets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。