arXiv:2606.01074cs.CL2026-06被引 1

0.1%压缩率下仍保持性能,组合降维与量化效果更优

When Is 0.1% Enough? Analyzing the Combined Effects of Dimensionality Reduction and Quantization on Text Embedding Compression

  • 联合使用降维与量化压缩文本嵌入
  • 部分场景可压缩至原大小0.1%且性能几乎不变
  • 压缩策略需根据任务类型选择

当前高性能文本嵌入模型通常输出高维实值向量,带来显著存储与计算开销。为缓解此问题,已有基于降维或量化的方法被提出;然而,二者联合使用的效应尚未充分研究。本文系统评估了结合降维与量化压缩文本嵌入的效果,覆盖四个MTEB任务类别及四款预训练嵌入模型。实验表明,联合使用降维与量化可实现远超单一方法的压缩效果,在某些设置下嵌入可缩减至原始大小的0.1%而性能几乎无损,且最优压缩策略依赖于具体任务。

原文摘要 · Abstract (English)

Recent high-performing text embedding models often output high-dimensional real-valued vectors, resulting in substantial storage and computational costs. To address this issue, compression methods based on dimensionality reduction or quantization have been proposed; however, the effects of combining dimensionality reduction and quantization have not been sufficiently investigated. In this paper, we systematically examine the effectiveness of compressing text embeddings by combining dimensionality reduction and quantization, using four MTEB task families and four pretrained embedding models. The experimental results demonstrate that combining dimensionality reduction and quantization enables substantially stronger compression than using either method alone, that in some settings embeddings can be reduced to as little as 0.1% of their original size with almost no performance degradation, and that the optimal compression strategy depends on the task.

嵌入压缩降维量化文本表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。