根据数据复杂度自动选择降维方法,省去试错成本。
Dataset-Adaptive Dimensionality Reduction
- 用结构复杂度指标衡量数据内在复杂性
- 可预测最优降维维度,避免无效尝试
- 适合需要高效降维的科研与工程场景
选择合适的降维(DR)方法及其最优超参数通常需大量试错,导致计算开销过大。为此,我们提出一种基于结构复杂度度量的数据集自适应降维优化方法。这些度量可量化数据集的内在复杂性,预测是否需要高维空间以准确表示数据。由于复杂数据在二维投影中常被误表示,利用这些度量可预测给定数据集上降维技术的最大可能准确率,从而消除冗余的优化尝试。我们设计并建立了这些结构复杂度度量的理论基础,并通过定量验证其对数据真实复杂度的有效逼近能力,确认其在指导数据集自适应降维流程中的适用性。最后,实验表明,我们的自适应流程显著提升了降维优化效率,且不牺牲准确性。
原文摘要 · Abstract (English)
Selecting the appropriate dimensionality reduction (DR) technique and determining its optimal hyperparameter settings that maximize the accuracy of the output projections typically involves extensive trial and error, often resulting in unnecessary computational overhead. To address this challenge, we propose a dataset-adaptive approach to DR optimization guided by structural complexity metrics. These metrics quantify the intrinsic complexity of a dataset, predicting whether higher-dimensional spaces are necessary to represent it accurately. Since complex datasets are often inaccurately represented in two-dimensional projections, leveraging these metrics enables us to predict the maximum achievable accuracy of DR techniques for a given dataset, eliminating redundant trials in optimizing DR. We introduce the design and theoretical foundations of these structural complexity metrics. We quantitatively verify that our metrics effectively approximate the ground truth complexity of datasets and confirm their suitability for guiding dataset-adaptive DR workflow. Finally, we empirically show that our dataset-adaptive workflow significantly enhances the efficiency of DR optimization without compromising accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。