arXiv:2604.21555cs.CL2026-04

用噪声和反义句测试嵌入模型对概念的稳定性,无需额外分类器。

Finding Meaning in Embeddings: Concept Separation Curves

论文配图:Finding Meaning in Embeddings: Concept Separation Curves
图 1 · 摘自论文原文
  • 通过引入语法噪声和语义否定,评估嵌入对概念的敏感度。
  • 概念分离曲线能清晰展示模型在不同句子长度下的概念区分能力。
  • 跨语言跨领域验证,适合研究嵌入质量的学者使用。

句子嵌入技术旨在将句子意义的关键概念编码到向量空间中。然而,现有评估方法多依赖额外分类器或下游任务,难以判断性能优劣是来自嵌入本身还是分类器行为。本文提出一种新型无分类器的评估方法,通过系统性地在句子中引入语法噪声和语义否定,量化其对嵌入结果的影响,并以概念分离曲线可视化模型对概念与表面变化的区分能力。该方法在荷兰语和英语多领域数据集上验证,覆盖不同句子长度,证明概念分离曲线具备可解释性、可复现性和跨模型适用性,能有效评估句子嵌入的概念稳定性。

原文摘要 · Abstract (English)

Sentence embedding techniques aim to encode key concepts of a sentence's meaning in a vector space. However, the majority of evaluation approaches for sentence embedding quality rely on the use of additional classifiers or downstream tasks. These additional components make it unclear whether good results stem from the embedding itself or from the classifier's behaviour. In this paper, we propose a novel method for evaluating the effectiveness of sentence embedding methods in capturing sentence-level concepts. Our approach is classifier-independent, allowing for an objective assessment of the model's performance. The approach adopted in this study involves the systematic introduction of syntactic noise and semantic negations into sentences, with the subsequent quantification of their relative effects on the resulting embeddings. The visualisation of these effects is facilitated by Concept Separation Curves, which show the model's capacity to differentiate between conceptual and surface-level variations. By leveraging data from multiple domains, employing both Dutch and English languages, and examining sentence lengths, this study offers a compelling demonstration that Concept Separation Curves provide an interpretable, reproducible, and cross-model approach for evaluating the conceptual stability of sentence embeddings.

嵌入评估概念分离无分类器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。