arXiv:2605.22202cs.CL2026-05

Embedding模型性能高低,可由其向量空间结构一致性预测。

Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance

论文配图:Structure Retention in Embedding Spaces as a Predictor of Benchmark Performance
图 1 · 摘自论文原文
  • 通过分析嵌入空间中近邻重叠与ICA幅值差异,发现其与任务表现强相关。
  • 25个模型在5个MTEB任务上测试,相关系数高达0.97。
  • 适用于研究嵌入机制、优化模型训练目标的研究者。

本文表明高性能嵌入模型在嵌入空间中具有高度一致的组织方式。我们在涵盖检索、双语挖掘、成对分类和摘要四个任务类别的五项MTEB任务上,评估了25个现代嵌入模型,包括英文和多语言场景。结果揭示:配对文本实例间的最近邻重叠度与独立成分分析(ICA)中的幅值差异,在不同任务上与模型表现存在强相关性(最高达0.97)。研究进一步发现,各类嵌入任务对线性程度和局部信息保留的依赖程度各异。该成果深化了对嵌入机制及其与模型性能关系的理解,并为未来训练目标设计及条件嵌入优化提供了启示。

原文摘要 · Abstract (English)

In this paper, we show that high-performing embedding models organize their embedding spaces in a consistent way. We evaluate 25 contemporary embedding models on five MTEB tasks spanning four diverse task categories (retrieval, bitext mining, pair classification, and summarization) in both English and multilingual settings, and reveal that nearest-neighbor overlap and magnitude differences in independent component analysis (ICA) between paired text instances strongly correlate (even up to 0.97) with performance on the given task. Ultimately, we show that embedding tasks display varying degrees of linearity and reliance on retention of local information. Our results further the understanding of embeddings, their relation to model performance, and shed light on possible future training objectives and optimizing conditional embeddings.

嵌入空间模型性能相关性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。