arXiv:2506.01435cs.CL2025-06ACL被引 17

高维提示嵌入冗余严重,压缩至1%仍保持高性能。

Redundancy, Isotropy, and Intrinsic Dimensionality of Prompt-based Text Embeddings

  • 对提示嵌入进行降维,仅保留25%维度性能几乎不变。
  • 分类与聚类任务降至原维度0.5%时仍保持低损耗。
  • 揭示了不同任务嵌入的内在维度与各向异性差异。

基于提示的文本嵌入模型在接收定制提示后生成特定任务嵌入,近期表现出优异性能。然而其嵌入维度常达数千,带来高存储与计算开销。本文研究后处理降维对分类、聚类、检索及语义文本相似性(STS)任务的影响。实验表明,仅保留25%原始维度的朴素降维导致性能轻微下降,说明嵌入高度冗余。尤其对于分类与聚类任务,嵌入压缩至不足原维度0.5%时,性能损失极小。通过内在维度与各向同性分析发现,分类与聚类嵌入具有较低内在维度和较弱各向同性,相较于检索与STS任务更易压缩。

原文摘要 · Abstract (English)

Prompt-based text embedding models, which generate task-specific embeddings upon receiving tailored prompts, have recently demonstrated remarkable performance. However, their resulting embeddings often have thousands of dimensions, leading to high storage costs and increased computational costs of embedding-based operations. In this paper, we investigate how post-hoc dimensionality reduction applied to the embeddings affects the performance of various tasks that leverage these embeddings, specifically classification, clustering, retrieval, and semantic textual similarity (STS) tasks. Our experiments show that even a naive dimensionality reduction, which keeps only the first 25% of the dimensions of the embeddings, results in a very slight performance degradation, indicating that these embeddings are highly redundant. Notably, for classification and clustering, even when embeddings are reduced to less than 0.5% of the original dimensionality the performance degradation is very small. To quantitatively analyze this redundancy, we perform an analysis based on the intrinsic dimensionality and isotropy of the embeddings. Our analysis reveals that embeddings for classification and clustering, which are considered to have very high dimensional redundancy, exhibit lower intrinsic dimensionality and less isotropy compared with those for retrieval and STS.

嵌入压缩冗余分析提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。