让模型自评自训,用评分标准提升推理能力
Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning
- 模型自己当考官,根据评分标准生成奖励信号进行强化学习
- 仅用4000条数据训练,就超过GPT-5在难题上的表现
- 适合想高效提升推理能力的开发者或研究者
开放式评估对大语言模型在真实场景中的部署至关重要。在研究HealthBench时发现,使用模型自身作为评分者并生成基于评分标准的奖励信号,能显著提升推理性能。令人惊讶的是,训练后的模型也变成了更强的评分者。受此启发,我们提出Self-Rewarding Rubric-Based Reinforcement Learning框架,实现更快、更高效的训练,且性能超越基线。在Qwen3-32B上,仅用HealthBench Easy的4000样本即可训练出在HealthBench Hard上表现优于GPT-5的模型。加入少量教师标注数据可进一步提升低能力模型的表现。
原文摘要 · Abstract (English)
Open-ended evaluation is essential for deploying large language models in real-world settings. In studying HealthBench, we observe that using the model itself as a grader and generating rubric-based reward signals substantially improves reasoning performance. Remarkably, the trained model also becomes a stronger grader. Motivated by this, we introduce Self-Rewarding Rubric-Based Reinforcement Learning for Open-Ended Reasoning, a lightweight framework that enables faster and more resource-efficient training while surpassing baselines. Remarkably, on Qwen3-32B, training with just the 4000-sample HealthBench Easy subset is sufficient to obtain a model that exceeds GPT-5 on HealthBench Hard. Incorporating a small amount of teacher-graded data further enhances performance for less capable models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。