arXiv:2411.15894cs.LGcs.AI2024-11NeurIPS被引 8

提出新方法提升参数化降维对局部结构的保留能力。

Navigating the Effect of Parametrization for Dimensionality Reduction

  • 引入硬负样本挖掘与强排斥力损失函数,改进参数化降维。
  • 在保持全局结构的同时,显著提升局部细节保留效果。
  • 适合关注降维中局部结构精度的研究者和工程师。

参数化降维方法因其可泛化至未见数据集的能力而日益受到重视,这正是传统方法所缺乏的优势。然而,从业者普遍存在一种误解,认为参数化与非参数化方法性能相当。本文揭示二者并不等价:参数化方法虽能保留全局结构,却会丢失大量局部细节。我们进一步证明,这类方法缺乏排斥负样本的能力,且损失函数的选择也影响性能。为此,我们提出了新方法 ParamRepulsor,融合硬负样本挖掘与强排斥力损失函数。该方法在不牺牲全局结构保真度的前提下,实现了参数化降维中局部结构保留的最先进水平。代码已开源:https://github.com/hyhuang00/ParamRepulsor。

原文摘要 · Abstract (English)

Parametric dimensionality reduction methods have gained prominence for their ability to generalize to unseen datasets, an advantage that traditional approaches typically lack. Despite their growing popularity, there remains a prevalent misconception among practitioners about the equivalence in performance between parametric and non-parametric methods. Here, we show that these methods are not equivalent -- parametric methods retain global structure but lose significant local details. To explain this, we provide evidence that parameterized approaches lack the ability to repulse negative pairs, and the choice of loss function also has an impact. Addressing these issues, we developed a new parametric method, ParamRepulsor, that incorporates Hard Negative Mining and a loss function that applies a strong repulsive force. This new method achieves state-of-the-art performance on local structure preservation for parametric methods without sacrificing the fidelity of global structural representation. Our code is available at https://github.com/hyhuang00/ParamRepulsor.

降维参数化结构保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。