arXiv:2601.08430cs.AI2026-01ACL被引 35

自动生成精细评分标准,提升大模型推理能力

RubricHub: A Comprehensive and Highly Discriminative Rubric Dataset via Automated Coarse-to-Fine Generation

  • 通过粗到精的自动化框架生成多维度评分标准
  • 构建超11万条目的跨领域数据集,显著提升评估精度
  • 适合需要高质量推理评测的AI研究者和开发者

强化学习结合可验证奖励(RLVR)在数学等推理密集型任务中取得显著进展。然而,开放生成任务因缺乏真实答案而难以优化。现有基于评分标准的评估方法存在扩展性差、标准粗糙的问题,导致监督上限。为此,我们提出一种自动化的粗到精评分生成框架,融合原则引导合成、多模型聚合与难度演化机制,生成全面且高度区分性的评估标准。基于该框架,我们构建了大规模(约11万条)、多领域的RubricHub数据集。通过两阶段后训练流程——基于评分的拒绝采样微调(RuFT)与强化学习(RuRL),实验表明,经后训练的Qwen3-14B在HealthBench上达到69.3分,超越如GPT-5等专有前沿模型,实现新SOTA。代码已开源。

原文摘要 · Abstract (English)

Reinforcement Learning with Verifiable Rewards (RLVR) has driven substantial progress in reasoning-intensive domains like mathematics. However, optimizing open-ended generation remains challenging due to the lack of ground truth. While rubric-based evaluation offers a structured proxy for verification, existing methods suffer from scalability bottlenecks and coarse criteria, resulting in a supervision ceiling effect. To address this, we propose an automated Coarse-to-Fine Rubric Generation framework. By synergizing principle-guided synthesis, multi-model aggregation, and difficulty evolution, our approach produces comprehensive and highly discriminative criteria capable of capturing the subtle nuances. Based on this framework, we introduce RubricHub, a large-scale ($\sim$110k) and multi-domain dataset. We validate its utility through a two-stage post-training pipeline comprising Rubric-based Rejection Sampling Fine-Tuning (RuFT) and Reinforcement Learning (RuRL). Experimental results demonstrate that RubricHub unlocks significant performance gains: our post-trained Qwen3-14B achieves state-of-the-art (SOTA) results on HealthBench (69.3), surpassing proprietary frontier models such as GPT-5. Our code is available at \href{https://github.com/teqkilla/RubricHub}{ this URL}.

评分标准强化学习大模型评测自动化生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。