arXiv:2502.20975cs.CL2025-02

用集合论评估句子嵌入的组合性,发现SBERT表现最佳。

Set-Theoretic Compositionality of Sentence Embeddings

  • 基于集合运算设计六项评估标准,检验嵌入的组合能力。
  • 7个经典模型和9个LLM模型中,SBERT在集合类组合上最优。
  • 新构建19.2万样本数据集,助力未来研究基准测试。

句子编码器在多种自然语言处理任务中起关键作用,因此准确评估其组合性质至关重要。然而,现有评估方法多聚焦于特定任务的表现,未能在任务无关背景下揭示句子嵌入的基本组合特性。本文借鉴经典集合论,提出基于三种核心“集合式”操作——文本重叠(TextOverlap)、文本差异(TextDifference)和文本并集(TextUnion)的六项评估准则。我们系统评估了7个经典模型和9个基于大语言模型(LLM)的句子编码器,结果表明,SBERT在各项指标上均表现出优异的集合类组合性,甚至优于最新大型语言模型。此外,我们构建了一个包含约19.2万个样本的新数据集,以支持未来对句子嵌入集合式组合性的基准研究。

原文摘要 · Abstract (English)

Sentence encoders play a pivotal role in various NLP tasks; hence, an accurate evaluation of their compositional properties is paramount. However, existing evaluation methods predominantly focus on goal task-specific performance. This leaves a significant gap in understanding how well sentence embeddings demonstrate fundamental compositional properties in a task-independent context. Leveraging classical set theory, we address this gap by proposing six criteria based on three core "set-like" compositions/operations: \textit{TextOverlap}, \textit{TextDifference}, and \textit{TextUnion}. We systematically evaluate $7$ classical and $9$ Large Language Model (LLM)-based sentence encoders to assess their alignment with these criteria. Our findings show that SBERT consistently demonstrates set-like compositional properties, surpassing even the latest LLMs. Additionally, we introduce a new dataset of ~$192$K samples designed to facilitate future benchmarking efforts on set-like compositionality of sentence embeddings.

句子嵌入集合论组合性评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。