arXiv:2603.10493cs.LG2026-03被引 2

提出一种无需假设的全新数据内在维度估计方法。

A Universal Nearest-Neighbor Estimator for Intrinsic Dimensionality

  • 基于最近邻距离比设计新估计算法,计算简单。
  • 理论证明该方法对任意数据分布均收敛到真实维度。
  • 适用于高维数据降维与结构分析,适合研究人员参考。

估计数据的内在维度(ID)是机器学习和计算机视觉中的基础问题,有助于理解高维观测背后的真正自由度。现有方法通常依赖几何或分布假设,在假设不成立时表现显著下降。本文提出一种基于最近邻距离比的新ID估计算法,计算简单且达到当前最优性能。更重要的是,我们提供了理论分析,证明该估计器具有‘通用性’——无论数据由何种分布生成,都能收敛到真实内在维度。我们在基准流形和真实数据集上进行了实验,验证了该方法的有效性。

原文摘要 · Abstract (English)

Estimating the intrinsic dimensionality (ID) of data is a fundamental problem in machine learning and computer vision, providing insight into the true degrees of freedom underlying high-dimensional observations. Existing methods often rely on geometric or distributional assumptions and can significantly fail when these assumptions are violated. In this paper, we introduce a novel ID estimator based on nearest-neighbor distance ratios that involves simple calculations and achieves state-of-the-art results. Most importantly, we provide a theoretical analysis proving that our estimator is \emph{universal}, namely, it converges to the true ID independently of the distribution generating the data. We present experimental results on benchmark manifolds and real-world datasets to demonstrate the performance of our estimator.

维度估计无假设机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。