无需标签即可稳定评估嵌入模型性能,尤其在高维空间表现优异。
FLARE: Task-agnostic embedding model evaluation through a normalization process

- 基于流模型的归一化流估计信息充分性,避开距离依赖的密度估计。
- 在11个数据集、8种嵌入模型上达0.90的斯皮尔曼相关系数,高维下仍稳定。
- 适合无标签场景下嵌入模型选型,尤其适用于高维数据任务。
当缺乏特定任务标签时,难以为特定目标语料选择合适的嵌入模型。现有基于核估计或高斯混合的无标签评估方法在高维空间中失效,导致排名不稳定。我们提出一种基于流的无标签嵌入评估方法(FLARE),利用归一化流直接从对数似然估计信息充分性,避免基于距离的密度估计。我们给出了有限样本边界,表明估计误差取决于数据流形的内在维度,而非原始嵌入维度。在11个数据集和8种嵌入模型上,FLARE在监督基准下达到0.90的斯皮尔曼等级相关系数,且在高维嵌入(d ≥ 3,584)中保持稳定,而现有无标签基线方法已崩溃。
原文摘要 · Abstract (English)
When task-specific labels are not available, it becomes difficult to select an embedding model for a specific target corpus. Existing labelless measures based on kernel estimators or Gaussian mixes fail in high-dimensional space, resulting in unstable rankings. We propose a flow-based labelless representation embedding evaluation (FLARE), which utilizes normalized streams to estimate information sufficiency directly from log-likelihood and avoid distance-based density estimation. We give a finite sample boundary, indicating that the estimation error depends on the intrinsic dimension of the data manifold rather than the original embedding dimension. On 11 datasets and 8 embedders, FLARE reached Spearman's $ρ$ of 0.90 under the supervised benchmark and remained stable in high-dimensional embeddings ($d \geq 3{,}584$) as the existing labelless baseline collapsed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。