arXiv:2601.05075cs.CL2026-01ACL被引 1

用语义偏好对齐提升大模型句向量,不损失生成能力

SemPA: Improving Sentence Embeddings of Large Language Models through Semantic Preference Alignment

  • 通过句子级直接偏好优化,对齐语义等价句子
  • 在多个任务上表现优于现有方法,且保持生成能力
  • 适合需要高质量句向量的下游应用

传统句向量方法在非生成式预训练模型上采用基于词元的对比学习。近年来出现基于生成式大语言模型(LLMs)的嵌入方法,但或依赖固定提示模板,或修改模型结构:前者无法进一步优化模型,后者损害生成能力。本文提出SemPA,一种通过语义偏好对齐增强句表示的同时保留LLM生成能力的新方法。利用句子级直接偏好优化(DPO)在重述生成任务上高效优化LLM,使模型学会区分语义等价句子,同时保持内在生成能力。理论上,我们在Plackett-Luce模型框架下建立了DPO与对比学习的正式关联。实验表明,SemPA在语义文本相似性任务及多个LLM基准测试中均取得更优的语义表示,且未牺牲模型的生成性能。

原文摘要 · Abstract (English)

Traditional sentence embedding methods employ token-level contrastive learning on non-generative pre-trained models. Recently, there have emerged embedding methods based on generative large language models (LLMs). These methods either rely on fixed prompt templates or involve modifications to the model architecture. The former lacks further optimization of the model and results in limited performance, while the latter alters the internal computational mechanisms of the model, thereby compromising its generative capabilities. We propose SemPA, a novel approach that boosts the sentence representations while preserving the generative ability of LLMs via semantic preference alignment. We leverage sentence-level Direct Preference Optimization (DPO) to efficiently optimize LLMs on a paraphrase generation task, where the model learns to discriminate semantically equivalent sentences while preserving inherent generative capacity. Theoretically, we establish a formal connection between DPO and contrastive learning under the Plackett-Luce model framework. Empirically, experimental results on both semantic textual similarity tasks and various benchmarks for LLMs show that SemPA achieves better semantic representations without sacrificing the inherent generation capability of LLMs.

句向量大模型偏好对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。