arXiv:2601.21192cs.AIcs.CL2026-01

用强化学习训练的推理模型,初始化嵌入模型效果不升反降。

Do Reasoning Models Enhance Embedding Models?

  • 通过层级相似性分析揭示推理模型改变局部几何但保留整体结构
  • 在MTEB和BRIGHT上验证:推理模型初始化无性能提升
  • 适合关注大模型语义表征机制的研究者阅读

当前最先进的嵌入模型多基于解码器仅有的大语言模型(LLM)骨干,经对比学习微调。随着基于可验证奖励强化学习(RLVR)训练的推理模型兴起,一个自然问题浮现:这些增强的推理能力是否能带来更优的语义表示?出乎意料的是,在MTEB和BRIGHT上的评估显示其效果为**零效应**:使用RLVR微调的骨干初始化嵌入模型,并未在相同训练流程下表现出一致优势。为解析这一悖论,我们提出层级表示相似性分析(HRSA)框架,从表示、几何和功能层面分解相似性。结果表明,尽管RLVR导致潜在流形的局部几何重构与坐标基漂移,但整体流形几何与线性读出能力得以保留。因此,后续对比学习使基础与推理初始化模型实现强对齐,我们称之为**流形重校准**。实证发现,不同于监督微调(SFT),RLVR是在既定语义空间内优化路径,而非根本性重塑该空间。

原文摘要 · Abstract (English)

State-of-the-art embedding models are increasingly derived from decoder-only Large Language Model (LLM) backbones adapted via contrastive learning. Given the emergence of reasoning models trained via Reinforcement Learning with Verifiable Rewards (RLVR), a natural question arises: do enhanced reasoning translate to superior semantic representations when these models serve as embedding initializations? Contrary to expectation, our evaluation on MTEB and BRIGHT reveals a **null effect**: embedding models initialized from RLVR-tuned backbones yield no consistent performance advantage over their base counterparts when subjected to identical training recipes. To unpack this paradox, we introduce **H**ierarchical **R**epresentation **S**imilarity **A**nalysis (HRSA), a framework that decomposes similarity across representation, geometry, and function levels. HRSA reveals that while RLVR induces irreversible latent manifold's local geometry reorganization and reversible coordinate basis drift, it preserves the global manifold geometry and linear readout. Consequently, subsequent contrastive learning drives strong alignment between base- and reasoning-initialized models, a phenomenon we term **Manifold Realignment**. Empirically, our findings suggest that unlike Supervised Fine-Tuning (SFT), RLVR optimizes trajectories within an existing semantic landscape rather than fundamentally restructuring the landscape itself.

嵌入模型强化学习语义表征流形分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。