综述数据内在维度估计方法,帮研究者选对工具。
A Survey of Dimension Estimation Methods
- 按几何特征分类:切线、参数、拓扑度量三类方法
- 发现多数方法在高维、噪声下表现不稳定,易过拟合
- 适合需要评估数据真实复杂度的研究者参考
高维数据通常具有内在结构,实际位于或靠近低维子集。理解数据的真实维度对把握其复杂性至关重要。已有大量维度估计方法被提出,但缺乏可靠使用指南。本文系统综述多种维度估计方法,按其利用的几何信息分为三类:检测局部线性结构的切线型方法、依赖维度相关概率分布的参数型方法,以及基于拓扑或度量不变量的方法。论文评估了这些方法在不同条件下的表现,包括对曲率和噪声的敏感性。重点分析了超参数选择鲁棒性、样本量需求、高维精度、估计精度及非线性几何适应性。在基准数据集上寻找最优超参数时普遍出现过拟合现象,表明许多方法泛化能力有限,难以推广到新数据集。
原文摘要 · Abstract (English)
It is a standard assumption that datasets in high dimension have an internal structure which means that they in fact lie on, or near, subsets of a lower dimension. In many instances it is important to understand the real dimension of the data, hence the complexity of the dataset at hand. A great variety of dimension estimators have been developed to find the intrinsic dimension of the data but there is little guidance on how to reliably use these estimators. This survey reviews a wide range of dimension estimation methods, categorising them by the geometric information they exploit: tangential estimators which detect a local affine structure; parametric estimators which rely on dimension-dependent probability distributions; and estimators which use topological or metric invariants. The paper evaluates the performance of these methods, as well as investigating varying responses to curvature and noise. Key issues addressed include robustness to hyperparameter selection, sample size requirements, accuracy in high dimensions, precision, and performance on non-linear geometries. In identifying the best hyperparameters for benchmark datasets, overfitting is frequent, indicating that many estimators may not generalise well beyond the datasets on which they have been tested.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。