arXiv:2507.09973cs.CLcs.AI2025-07

4亿参数小模型媲美175倍大的语言模型,高效实现推理偏好建模。

Tiny Reward Models

  • 用4亿参数的双向掩码模型,结合提示工程与轻量微调技术。
  • 在RewardBench上表现接近超大模型,推理成本降低95%以上。
  • 适合资源受限场景下的安全与推理偏好建模任务。

基于解码器的大语言模型已成为强化学习中人类反馈(RLHF)的主流奖励建模架构。然而,随着奖励模型在测试阶段日益广泛应用,其推理开销成为关注焦点。我们提出TinyRM,一系列仅含4亿参数的小型双向掩码语言模型(MLMs),在推理与安全偏好建模任务上可媲美参数量超过175倍的大型模型。TinyRM结合了FLAN风格提示、方向性低秩适配(DoRA)与层冻结策略,在RewardBench上取得优异表现,同时显著降低资源消耗。实验表明,小模型在特定领域任务中受益于定制化微调策略,尤其在推理任务中,轻量化微调方法尤为有效。尽管通用模型与对话偏好建模仍面临挑战,初步结果凸显了轻量级双向架构在偏好建模中的高效性与可扩展性。

原文摘要 · Abstract (English)

Large decoder-based language models have become the dominant architecture for reward modeling in reinforcement learning from human feedback (RLHF). However, as reward models are increasingly deployed in test-time strategies, their inference costs become a growing concern. We present TinyRM, a family of small, bidirectional masked language models (MLMs) with as few as 400 million parameters, that rival the capabilities of models over 175 times larger on reasoning and safety preference modeling tasks. TinyRM combines FLAN-style prompting, Directional Low-Rank Adaptation (DoRA), and layer freezing to achieve strong performance on RewardBench, despite using significantly fewer resources. Our experiments suggest that small models benefit from domain-specific tuning strategies, particularly in reasoning, where lightweight finetuning methods are especially effective. While challenges remain in building generalist models and conversational preference modeling, our preliminary results highlight the promise of lightweight bidirectional architectures as efficient, scalable alternatives for preference modeling.

奖励模型轻量模型微调推理效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。