arXiv:2504.16411cs.CL2025-04被引 5

无需微调,用大模型生成条件化文本嵌入

Out-of-the-Box Conditional Text Embeddings from Large Language Models

  • 基于因果大模型和条件提示,无监督生成嵌入
  • 在语义相似度与聚类任务上媲美有监督方法
  • 可解释性强,支持词生成分析与嵌入可视化

条件化文本嵌入是一种捕捉文本在特定视角下语义变化的表示方法。以往方法依赖大量标注数据微调模型,成本高昂。本文提出 PonTE,一种无需微调的无监督条件化文本嵌入方法,利用因果大语言模型与条件提示生成嵌入。在条件语义文本相似度与文本聚类任务上的实验表明,PonTE 能生成有效嵌入,性能接近有监督方法。此外,通过分析提示后的词生成过程与嵌入可视化,验证了其可解释性。

原文摘要 · Abstract (English)

Conditional text embedding is a proposed representation that captures the shift in perspective on texts when conditioned on a specific aspect. Previous methods have relied on extensive training data for fine-tuning models, leading to challenges in terms of labor and resource costs. We propose PonTE, a novel unsupervised conditional text embedding method that leverages a causal large language model and a conditional prompt. Through experiments on conditional semantic text similarity and text clustering, we demonstrate that PonTE can generate useful conditional text embeddings and achieve performance comparable to supervised methods without fine-tuning. We also show the interpretability of text embeddings with PonTE by analyzing word generation following prompts and embedding visualization.

文本嵌入无监督学习可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。