arXiv:2511.15256cs.LGcs.CV2025-11被引 2

用GRPO强化学习优化表示模型,提升下游任务表现

GRPO-RM: Fine-Tuning Representation Models via GRPO-Driven Reinforcement Learning

  • 用预定义输出集替代采样,实现表示模型的组相对优化
  • 在多个真实数据集上显著提升表示模型性能
  • 适合需要微调表示模型的研究者和工程师

Group Relative Policy Optimization (GRPO) 是一种用于微调大语言模型的强化学习方法,在 DeepSeek-R1 等实际应用中表现优异。本文探讨了 GRPO 是否可推广至表示学习模型。为此,我们提出 GRPO-RM,研究类 GRPO 策略在后训练表示模型中的表现。方法通过构建预定义输出集,替代语言模型中的词元序列采样,生成输出组,以支持 GRPO 的概率驱动优化。同时设计专用奖励函数,适配表示模型的特性。在多个真实数据集上的大量实验验证了该方法的有效性。

原文摘要 · Abstract (English)

The Group Relative Policy Optimization (GRPO), a reinforcement learning method used to fine-tune large language models (LLMs), has proved its effectiveness in practical applications such as DeepSeek-R1. It raises a question whether GRPO can be generalized to representation learning models. In this paper, we propose Group Relative Policy Optimization for Representation Model (GRPO-RM), and investigate the performance of GRPO-like policy in post-training representation models. Specifically, our method establishes a predefined output set to functionally replace token sequence sampling in LLMs, thereby generating an output group, which is essential for the probability-driven optimization of GRPO. In addition, a specialized reward function is designed to accommodate the properties of representation models. Extensive experiments are conducted on various real-world datasets to validate the effectiveness of our proposed method.

强化学习表示学习微调模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。