arXiv:2604.23573stat.MLcs.LG2026-04被引 1

用费马距离提升高维半监督分类,让无标签数据更有效。

High-dimensional Semi-supervised Classification via the Fermat Distance

论文配图:High-dimensional Semi-supervised Classification via the Fermat Distance
图 1 · 摘自论文原文
  • 基于费马距离设计加权k近邻和多维缩放分类器。
  • 理论证明其误差随样本量指数下降,优于现有方法。
  • 适合标签少、数据高维的场景,如生物医学分析。

半监督分类在标签数据稀缺而无标签数据庞大的情况下广泛应用。针对高维数据,本文基于流形与聚类假设,利用密度敏感的费马距离,提出加权k近邻(k-NN)分类器和多维缩放(MDS)诱导的分类器。利用大目标维度的MDS,使线性分类器能有效处理复杂流形数据。理论上,我们推导出簇内期望过失风险的紧下界,并证明使用真实费马距离的加权k-NN分类器为极小极大最优。此外,我们明确量化了无标签数据的价值:估计费马距离带来的误差随合并样本量呈指数衰减,该速率远快于文献中相关结果。在合成与真实数据集上的大量实验表明,本方法性能优于或媲美当前最先进的图基半监督分类器。

原文摘要 · Abstract (English)

Semi-supervised classification, where unlabeled data are massive but labeled data are limited, often arises in machine learning applications. We address this challenge under high-dimensional data by leveraging the manifold and cluster assumptions. Based on the Fermat distance, a density-sensitive metric that naturally encodes the cluster assumption, we propose the weighted $k$-nearest neighbors (NN) classifier and multidimensional scaling (MDS)-induced classifiers. The use of MDS with a large target dimension allows the effective application of linear classifiers to complex manifold data. Theoretically, we derive a sharp lower bound for the expected excess risk within clusters and prove that the weighted $k$-NN classifier utilizing the true Fermat distance is minimax optimal. Furthermore, we explicitly quantify the utility of unlabeled data by showing that the error arising from estimating the Fermat distance decays exponentially with the pooled sample size. Such a rate is much faster than the related rates in the literature. Extensive experiments on synthetic and real datasets demonstrate competitive or superior performance of our approaches compared to state-of-the-art graph-based semi-supervised classifiers.

半监督学习高维数据费马距离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。