arXiv:2510.10224cs.CLcs.IR2025-10

通过预测关键词构建无监督文本表示,提升模型对语义核心的理解能力。

Text2Token: Unsupervised Text Representation Learning with Token Target Prediction

  • 基于关键词预测任务设计无监督学习框架
  • 在MTEB v2上性能媲美顶尖对比学习模型
  • 适合研究无监督表示学习与词汇空间优化的学者

无监督文本表示学习(TRL)是自然语言处理中的基础任务,有助于利用网络未标注文本提升搜索与推荐效果。近期实证研究表明,高质量表示与输入文本的关键词对齐,揭示了表示空间与词表空间之间的潜在关联。受此启发,我们重新审视生成任务,提出一种基于关键词预测的无监督生成框架Text2Token。该框架利用精心构建的目标词分布作为监督信号。为生成高质量目标词分布,我们分析了先进嵌入模型的词对齐特性,识别出两类关键词:(1)文本中的有意义词;(2)超越文本的语义衍生词。基于此,我们提出两种方法——数据驱动和模型衍生——从数据或LLM主干中生成合成关键词目标。在MTEB v2基准上的实验表明,Text2Token性能可与最先进的无监督对比学习模型LLM2Vec相媲美。进一步分析显示,词表空间与表示空间在训练过程中协同优化并趋向最优解,为未来研究提供了新思路。

原文摘要 · Abstract (English)

Unsupervised text representation learning (TRL) is a fundamental task in natural language processing, which is beneficial for improving search and recommendations with the web's unlabeled texts. A recent empirical study finds that the high-quality representation aligns with the key token of the input text, uncovering the potential connection between representation space and vocabulary space. Inspired by the findings, we revisit the generative tasks and develop an unsupervised generative framework for TRL, Text2Token. The framework is based on the token target prediction task, utilizing carefully constructed target token distribution as supervisory signals. To construct the high-quality target token distribution, we analyze the token-alignment properties with advanced embedders and identify two essential categories of key tokens: (1) the meaningful tokens in the text and (2) semantically derived tokens beyond the text. Based on these insights, we propose two methods -- data-driven and model-derived -- to construct synthetic token targets from data or the LLM backbone. Experiments on the MTEB v2 benchmark demonstrate that Text2Token achieves performance competitive with the state-of-the-art embedder with unsupervised contrastive learning, LLM2Vec. Our analysis further shows that vocabulary and representation spaces optimize together and toward the optimum solution during training, providing new ideas and insights for future work.

无监督学习文本表示关键词预测词表对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。