arXiv:2509.24291cs.CLcs.AI2025-09被引 8

用生成式迭代优化让大模型更懂语义,效果超越传统方法。

Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement

  • 通过自回归生成软标记,逐步优化句子嵌入表示。
  • 在MTEB基准上优于现有基于LLM的嵌入方法,且推理时增加生成长度能持续提升效果。
  • 适合关注语义表征与大模型生成能力结合的研究者。

现有的基于大语言模型(LLM)的嵌入方法通常采用编码器单一范式,将LLM视为静态特征提取器,忽略了其核心生成能力。我们提出GIRCSE(生成式迭代精炼对比句子嵌入),一种利用自回归生成迭代优化语义表示的新框架。通过在对比目标下生成一系列优化的软标记,GIRCSE捕捉到编码器方法常遗漏的潜在概念和隐含语义。为此,我们设计了迭代对比精炼(ICR)目标,促使每一步精炼都产生更优表示。大量实验表明,GIRCSE在MTEB基准和指令遵循任务中优于强基线。此外,GIRCSE展现出涌现的测试时缩放特性:推理时生成更多标记可稳定提升嵌入质量。结果确立了生成式迭代精炼作为表征学习的新范式。

原文摘要 · Abstract (English)

Existing large language model (LLM)-based embeddings typically adopt an encoder-only paradigm, treating LLMs as static feature extractors and overlooking their core generative strengths. We introduce GIRCSE (Generative Iterative Refinement for Contrastive Sentence Embeddings), a novel framework that leverages autoregressive generation to iteratively refine semantic representations. By producing sequences of soft tokens optimized under contrastive objective, GIRCSE captures latent concepts and implicit semantics that encoder-only methods often miss. To guide this process, we propose an Iterative Contrastive Refinement (ICR) objective that encourages each refinement step to yield better representations. Extensive experiments show that GIRCSE outperforms strong LLM-based embedding baselines on the MTEB benchmark and instruction-following tasks. Moreover, GIRCSE exhibits an emergent test-time scaling property: generating more tokens at inference steadily improves embedding quality. Our results establish generative iterative refinement as a new paradigm for representation learning.

生成嵌入语义表征LLM应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。