实时分析两组流数据的相关性,适合高维在线场景。
Sliding Window Informative Canonical Correlation Analysis
- 用滑动窗口+在线PCA实时估算CCA特征。
- 在高维数据上表现稳定,理论性能有保障。
- 适合金融、传感器等实时数据监测场景。
典型相关分析(CCA)用于发现两组数据之间的相关特征集。本文提出一种新的CCA扩展方法——滑动窗口有信息的典型相关分析(SWICCA),适用于在线流数据场景。该方法以在线主成分分析(PCA)算法为后端,结合少量滑动窗口样本,实现实时估计CCA成分。文中阐述了算法动机与设计,通过数值模拟评估其性能,并提供了理论性能保证。SWICCA方法适用于极高维度数据,且具备可扩展性,文中还展示了一个真实数据案例,验证其实际应用能力。
原文摘要 · Abstract (English)
Canonical correlation analysis (CCA) is a technique for finding correlated sets of features between two datasets. In this paper, we propose a novel extension of CCA to the online, streaming data setting: Sliding Window Informative Canonical Correlation Analysis (SWICCA). Our method uses a streaming principal component analysis (PCA) algorithm as a backend and uses these outputs combined with a small sliding window of samples to estimate the CCA components in real time. We motivate and describe our algorithm, provide numerical simulations to characterize its performance, and provide a theoretical performance guarantee. The SWICCA method is applicable and scalable to extremely high dimensions, and we provide a real-data example that demonstrates this capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。