通过球面约束提升嵌入分散性,改善图像与文本任务性能
Keep your distance: learning dispersed embeddings on $\mathbb{S}_m$
- 用MMD思想重新定义向量间分散度,增强理论可解释性
- 提出在线版Lloyd算法,高效实现高维空间的分散嵌入
- 直接利用球面几何特性设计新方法,适用于多场景任务
在高维空间中学习充分分离的特征(如文本或图像嵌入)对众多机器学习应用至关重要。通过使无关向量尽可能远离,可有效实现特征分离。将特征约束在超球面上,可将分散性问题与数学和物理中的经典问题关联,但这些理论解仅适用于低维有限情形。在表示学习中,通常需处理大量高维特征,且分散性常与任务目标权衡,导致现有理论和数值方法难以适用。因此,普遍采用基于梯度的方法,通过最小化成对距离函数来促进分散。本文首先综述来自不同领域的现有方法,建立新联系并揭示共性;接着提出基于最大均值差异(MMD)动机的成对分散重释;引入著名的K-Means Lloyd算法的在线变体,作为通用域上的有效正则化器;最后提出一种直接利用超球面性质的新分散方法。实验表明,分散性在图像分类和自然语言处理中至关重要,不同算法在不同场景下呈现差异化权衡。
原文摘要 · Abstract (English)
Learning well-separated features in high-dimensional spaces, such as text or image embeddings, is crucial for many machine learning applications. Achieving such separation can be effectively accomplished through the dispersion of embeddings, where unrelated vectors are pushed apart as much as possible. By constraining features to be on a hypersphere, we can connect dispersion to well-studied problems in mathematics and physics, where optimal solutions are known for limited low-dimensional cases. However, in representation learning we typically deal with a large number of features in high-dimensional space, and moreover, dispersion is usually traded off with some other task-oriented training objective, making existing theoretical and numerical solutions inapplicable. Therefore, it is common to rely on gradient-based methods to encourage dispersion, usually by minimizing some function of the pairwise distances. In this work, we first give an overview of existing methods from disconnected literature, making new connections and highlighting similarities. Next, we introduce some new angles. We propose to reinterpret pairwise dispersion using a maximum mean discrepancy (MMD) motivation. We then propose an online variant of the celebrated Lloyd's algorithm, of K-Means fame, as an effective alternative regularizer for dispersion on generic domains. Finally, we derive a novel dispersion method that directly exploits properties of the hypersphere. Our experiments show the importance of dispersion in image classification and natural language processing tasks, and how algorithms exhibit different trade-offs in different regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。