用大模型摘要训练科学文本语义嵌入,提升相似性判断准确性。
SemCSE: Semantic Contrastive Sentence Embeddings Using LLM-Generated Summaries For Scientific Abstracts
- 用大模型生成摘要作为对比学习信号,强化语义表征。
- 在新基准上实现更强的语义区分能力,优于传统引用方法。
- 适合需要精准理解科学文本语义的研究者使用。
我们提出SemCSE,一种无监督学习科学文本语义嵌入的方法。基于对比学习的最新进展,该方法利用大语言模型(LLM)生成的科学摘要,训练模型使语义相关的摘要在嵌入空间中更接近。这种目标确保模型捕捉文本的真实语义内容,不同于依赖引用的传统方法,后者未必反映语义相似性。为此,我们设计了一个新基准,用于评估模型对科学文本语义内容的理解与编码能力,结果表明我们的方法在嵌入空间中实现了更强的语义分离。此外,在全面的SciRepEval基准测试中,SemCSE在同规模模型中达到最先进性能,凸显了语义导向训练的优势。
原文摘要 · Abstract (English)
We introduce SemCSE, an unsupervised method for learning semantic embeddings of scientific texts. Building on recent advances in contrastive learning for text embeddings, our approach leverages LLM-generated summaries of scientific abstracts to train a model that positions semantically related summaries closer together in the embedding space. This resulting objective ensures that the model captures the true semantic content of a text, in contrast to traditional citation-based approaches that do not necessarily reflect semantic similarity. To validate this, we propose a novel benchmark designed to assess a model's ability to understand and encode the semantic content of scientific texts, demonstrating that our method enforces a stronger semantic separation within the embedding space. Additionally, we evaluate SemCSE on the comprehensive SciRepEval benchmark for scientific text embeddings, where it achieves state-of-the-art performance among models of its size, thus highlighting the benefits of a semantically focused training approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。