arXiv:2608.25887stat.MLcs.LG2026-08

通过近邻关系高效估算高信息投影,提升数据降维效果。

Efficient Estimation of High Information Projections using Nearest Neighbours

论文配图:Efficient Estimation of High Information Projections using Nearest Neighbours
图 1 · 摘自论文原文
  • 基于近邻对局部协方差建模,构造密度信息矩阵的谱分解。
  • 在标准条件下一致估计密度信息矩阵,理论保障强。
  • 适用于聚类与异常检测,计算高效,适合实际应用。

提出一种直观且高效的降维方法,用于发现多变量数据中有趣的投影方向。该方法基于增强数据中最近邻关系的思想,通过构造一个编码局部协方差结构的矩阵,其局部协方差由每个点的最近邻对捕捉。在标准正则条件下,该矩阵是所谓“密度信息矩阵”(DIM)的一致估计器,而DIM是非参数化的费舍尔信息矩阵的对应物。已有研究表明,DIM的谱分解与独立成分分析及监督场景下的充分降维问题密切相关。然而,现有DIM估计方法计算成本高,且仅针对与真实密度平方成比例的代理密度。此外,本文还探讨了该方法在聚类分析和异常检测等下游任务中的实际效用。

原文摘要 · Abstract (English)

An intuitive method for dimensionality reduction is proposed, which is highly effective for finding interesting projections of multivariate data. Following similar intuitive motivation to a number of existing techniques, the proposed method is based on enhancing the nearest neighbour relationships in the data. The proposed projection arises from the spectral decomposition of a matrix designed to encode the local covariance structure in the data, where the local covariance at a point is captured by pairs of its nearest neighbours. We show that under standard regularity conditions this matrix is a consistent estimator of the so-called ``Density Information Matrix'' (DIM); a non-parametric analogue of the Fisher Information Matrix. Spectral decompositions of DIMs have been shown to be connected with the important problems of Independent Components Analysis and, in the supervised context, Sufficient Dimension Reduction. However, existing estimators of the DIM are computationally expensive to compute and only target the DIM of a surrogate density, which is proportional to the square of the true underlying density. In addition, we go on to explore the practical utility of our method in aiding the downstream tasks of cluster analysis and outlier detection.

降维近邻密度估计信息矩阵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。