arXiv:2505.12540cs.LG2025-05NeurIPS被引 64

无需配对数据即可跨模型翻译文本嵌入,保留语义结构。

Harnessing the Universal Geometry of Embeddings

  • 通过无监督方法将嵌入映射到通用潜在空间
  • 不同架构/参数/数据训练的模型间相似度高达0.92
  • 适用于嵌入安全分析,适合关注向量数据库风险的研究者

我们提出首个无需成对数据、编码器或预定义匹配集的文本嵌入跨空间翻译方法。该无监督方法可将任意嵌入映射至一个通用潜在表示(即柏拉图表征假说所预测的通用语义结构)。翻译结果在不同架构、参数量和训练数据的模型对之间实现了高余弦相似度(最高达0.92)。将未知嵌入转换至另一空间的同时保持其几何结构,对向量数据库安全具有重大影响:攻击者仅需访问嵌入向量,即可提取足够信息进行文档分类与属性推断。

原文摘要 · Abstract (English)

We introduce the first method for translating text embeddings from one vector space to another without any paired data, encoders, or predefined sets of matches. Our unsupervised approach translates any embedding to and from a universal latent representation (i.e., a universal semantic structure conjectured by the Platonic Representation Hypothesis). Our translations achieve high cosine similarity across model pairs with different architectures, parameter counts, and training datasets. The ability to translate unknown embeddings into a different space while preserving their geometry has serious implications for the security of vector databases. An adversary with access only to embedding vectors can extract sensitive information about the underlying documents, sufficient for classification and attribute inference.

嵌入翻译语义结构向量安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。