arXiv:2509.01198cs.LGcs.AI2025-09

提出关系保持损失,让降维后向量仍保正交与线性独立性。

Preserving Vector Space Properties in Dimensionality Reduction: A Relationship Preserving Loss Framework

  • 设计关系保持损失,通过对比高维与低维的格拉姆矩阵差异来优化投影。
  • 实验表明降维后下游任务性能基本不变,证明关键向量空间性质得以保留。
  • 适用于跨域对齐、知识蒸馏、联邦学习等需几何一致性场景。

降维可能破坏向量空间中的正交性和线性独立性,这些性质对跨模态检索、聚类和分类任务至关重要。本文提出关系保持损失(RPL),一种通过最小化高维数据与低维嵌入间关系矩阵(如格拉姆或余弦矩阵)差异来保持这些性质的损失函数。RPL 可用于训练非线性投影的神经网络,并基于矩阵扰动理论提供误差界支持。初步实验表明,使用 RPL 能有效降低嵌入维度,同时在下游任务中保持较高性能,原因在于关键向量空间性质得以保留。尽管本文聚焦于降维应用,但该损失也可推广至跨域对齐、迁移学习、知识蒸馏、公平性与不变性、去中心化、图与流形学习及联邦学习等场景,其中分布式嵌入需维持几何一致性。

原文摘要 · Abstract (English)

Dimensionality reduction can distort vector space properties such as orthogonality and linear independence, which are critical for tasks including cross-modal retrieval, clustering, and classification. We propose a Relationship Preserving Loss (RPL), a loss function that preserves these properties by minimizing discrepancies between relationship matrices (e.g., Gram or cosine) of high-dimensional data and their low-dimensional embeddings. RPL trains neural networks for non-linear projections and is supported by error bounds derived from matrix perturbation theory. Initial experiments suggest that RPL reduces embedding dimensions while largely retaining performance on downstream tasks, likely due to its preservation of key vector space properties. While we describe here the use of RPL in dimensionality reduction, this loss can also be applied more broadly, for example to cross-domain alignment and transfer learning, knowledge distillation, fairness and invariance, dehubbing, graph and manifold learning, and federated learning, where distributed embeddings must remain geometrically consistent.

降维向量空间损失函数嵌入

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。