arXiv:2503.09485cs.LGstat.ML2025-03被引 1

提出一种高效估算数据内在维度的新方法,降低高维数据处理成本。

A Novel Approach for Intrinsic Dimension Estimation

  • 仅通过矩阵-向量乘法实现内在维度估计,计算高效。
  • 在真实数据集上表现优于现有主流方法,稳定性强。
  • 适合大规模数据的降维前评估,尤其适用于资源受限场景。

真实数据由于其固有特性具有复杂的非线性结构,这些非线性与高维特征常导致空隙现象及著名的维数灾难问题。在低维空间中寻找数据的近似最优表示(即降维)是提升机器学习任务成功率的有效手段。然而,估算达到近似最优表示所需的最低数据维度(内在维度)往往代价高昂,尤其是在处理大数据时。本文提出一种高效且鲁棒的内在维度估计方法,该方法仅依赖于矩阵-向量乘积,适用于各类降维算法。通过实验对比,验证了所提方法在多个基准数据集上的性能优越性,显著提升了估计效率与稳定性。

原文摘要 · Abstract (English)

The real-life data have a complex and non-linear structure due to their nature. These non-linearities and the large number of features can usually cause problems such as the empty-space phenomenon and the well-known curse of dimensionality. Finding the nearly optimal representation of the dataset in a lower-dimensional space (i.e. dimensionality reduction) offers an applicable mechanism for improving the success of machine learning tasks. However, estimating the required data dimension for the nearly optimal representation (intrinsic dimension) can be very costly, particularly if one deals with big data. We propose a highly efficient and robust intrinsic dimension estimation approach that only relies on matrix-vector products for dimensionality reduction methods. An experimental study is also conducted to compare the performance of proposed method with state of the art approaches.

降维维度估计高效算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。