arXiv:2507.13094math.OCcs.LG2025-07被引 2

无监督学习最优特征距离,提升分类性能。

Unsupervised Ground Metric Learning

  • 提出随机函数迭代算法,证明线性收敛
  • 验证最优传输距离可被马氏距离替代
  • 引入图拉普拉斯方法,简化为线性问题

无标签数据的分类仍是挑战性问题,通常依赖于特征间合适的距离度量,即度量学习。近期,Huizing、Cantini 和 Peyré 提出同时学习样本与特征间的最优传输(OT)成本矩阵。这转化为寻找某非线性映射生成的OT距离的正特征向量。本文从算法与建模两方面研究无监督度量学习。首先,分析合适算法及其收敛性,提出使用随机函数迭代算法,并证明其在设定下线性收敛,尽管算子不满足此前要求的拟压缩性。其次,探讨能否用其他距离替代OT距离,发现马氏类距离可融入该框架。进一步,基于图拉普拉斯的方法仅需处理线性函数,可直接应用线性代数算法。

原文摘要 · Abstract (English)

Data classification without access to labeled samples remains a challenging problem. It usually depends on an appropriately chosen distance between features, a topic addressed in metric learning. Recently, Huizing, Cantini and Peyré proposed to simultaneously learn optimal transport (OT) cost matrices between samples and features of the dataset. This leads to the task of finding positive eigenvectors of a certain nonlinear function that maps cost matrices to OT distances. Having this basic idea in mind, we consider both the algorithmic and the modeling part of unsupervised metric learning. First, we examine appropriate algorithms and their convergence. In particular, we propose to use the stochastic random function iteration algorithm and prove that it converges linearly for our setting, although our operators are not paracontractive as it was required for convergence so far. Second, we ask the natural question if the OT distance can be replaced by other distances. We show how Mahalanobis-like distances fit into our considerations. Further, we examine an approach via graph Laplacians. In contrast to the previous settings, we have just to deal with linear functions in the wanted matrices here, so that simple algorithms from linear algebra can be applied.

度量学习最优传输无监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。