arXiv:2507.20783cs.CL2025-07综述被引 7

剖析预训练语言模型在通用文本嵌入中的核心作用

On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey

  • 以预训练模型提取嵌入,通过对比学习优化表示
  • 支持多语言、多模态及代码理解等高级功能
  • 适合想了解文本嵌入发展脉络的研究者

文本嵌入因在检索、分类、聚类、双语挖掘和摘要等众多自然语言处理任务中的有效性而备受关注。随着预训练语言模型(PLMs)的出现,通用文本嵌入(GPTE)因其能够生成丰富且可迁移的表示而广受欢迎。GPTE 的典型架构通常利用 PLMs 提取密集文本表示,并在大规模成对数据集上通过对比学习进行优化。本文综述了 PLMs 时代下 GPTE 的发展,重点分析了 PLMs 在其中扮演的角色。首先,我们探讨其基础架构,阐述 PLMs 在嵌入提取、表达能力增强、训练策略、学习目标和数据构建中的基本作用。随后,介绍由 PLMs 支持的进阶功能,包括多语言支持、多模态融合、代码理解及场景自适应。最后,指出未来可能的研究方向,如排序融合、安全性考量、偏见缓解、结构信息融入以及嵌入的认知扩展。本综述旨在为新学者和资深研究者提供当前状态与未来潜力的参考。

原文摘要 · Abstract (English)

Text embeddings have attracted growing interest due to their effectiveness across a wide range of natural language processing (NLP) tasks, including retrieval, classification, clustering, bitext mining, and summarization. With the emergence of pretrained language models (PLMs), general-purpose text embeddings (GPTE) have gained significant traction for their ability to produce rich, transferable representations. The general architecture of GPTE typically leverages PLMs to derive dense text representations, which are then optimized through contrastive learning on large-scale pairwise datasets. In this survey, we provide a comprehensive overview of GPTE in the era of PLMs, focusing on the roles PLMs play in driving its development. We first examine the fundamental architecture and describe the basic roles of PLMs in GPTE, i.e., embedding extraction, expressivity enhancement, training strategies, learning objectives, and data construction. We then describe advanced roles enabled by PLMs, including multilingual support, multimodal integration, code understanding, and scenario-specific adaptation. Finally, we highlight potential future research directions that move beyond traditional improvement goals, including ranking integration, safety considerations, bias mitigation, structural information incorporation, and the cognitive extension of embeddings. This survey aims to serve as a valuable reference for both newcomers and established researchers seeking to understand the current state and future potential of GPTE.

文本嵌入预训练模型综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。