通过联合与交叉协方差提升低采样数据中的信号可检测性
Better Together: Cross and Joint Covariances Enhance Signal Detectability in Undersampled Data
- 构建联合、交叉与自协方差矩阵,对比信号检测性能
- 联合与交叉协方差在样本不足时更早恢复共享信号
- 维度不匹配程度决定联合或交叉方法更优,适用于线性相关分析
许多数据科学任务需要从两个高维变量中检测共同信号。基于随机矩阵理论,我们确定了在采样噪声引起的相关性背景中,何时能从样本相关性中检测并重构该信号。研究了三种由两个高维变量构造的协方差矩阵:各自的自协方差、交叉协方差,以及拼接后联合变量的自协方差(包含自相关和交叉相关块)。所有矩阵均呈现预期的 Baik-Ben Arous-Péché 可检测相变。结果表明,联合与交叉协方差矩阵始终比自协方差更早重建共享信号。联合与交叉方法的优劣取决于两变量维度的失配程度。这些发现为选择检测线性相关性的合适方法提供了依据,并可能推广至非线性统计依赖关系。
原文摘要 · Abstract (English)
Many data-science applications involve detecting a shared signal between two high-dimensional variables. Using random matrix theory methods, we determine when such signal can be detected and reconstructed from sample correlations, despite the background of sampling noise induced correlations. We consider three different covariance matrices constructed from two high-dimensional variables: their individual self covariance, their cross covariance, and the self covariance of the concatenated (joint) variable, which incorporates the self and the cross correlation blocks. We observe the expected Baik, Ben Arous, and Péché detectability phase transition in all these covariance matrices, and we show that joint and cross covariance matrices always reconstruct the shared signal earlier than the self covariances. Whether the joint or the cross approach is better depends on the mismatch of dimensionalities between the variables. We discuss what these observations mean for choosing the right method for detecting linear correlations in data and how these findings may generalize to nonlinear statistical dependencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。