arXiv:2605.07461cs.CL2026-05被引 3

让评分标准内化为模型推理的指导,提升生成质量。

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance

论文配图:Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance
图 1 · 摘自论文原文
  • 将评分标准融入推理过程,作为内部生成引导。
  • 在多个基准上平均性能超越基线3.87分。
  • 适合需要高质量开放性回答的场景使用。

评分标准广泛用于评估无法验证的开放式任务,近期研究将其引入强化学习的奖励系统。然而,现有框架通常将评分标准视为与策略推理过程分离的外部评价工具,使其仅能事后测量,无法主动引导生成。本文提出Think-with-Rubrics,一种新型指令遵循任务范式。该方法将评分标准生成整合进推理上下文,使评分标准从独立产物转变为大语言模型生成的内部指导。训练时,模型按顺序生成评分标准和回答,由已训练的评分标准验证器联合监督答案与自生成或黄金评分标准之间的一致性。多个基准上的实验表明,Think-with-Rubrics在平均性能上比由黄金评分标准监督的基于评分标准的奖励基线高出3.87分。实验还揭示了其机制:黄金评分标准监督提升自生成评分标准的质量,而自生成评分标准监督则增强回答的内部一致性。

原文摘要 · Abstract (English)

Rubrics have been extensively utilized for evaluating unverifiable, open-ended tasks, with recent research incorporating them into reward systems for reinforcement learning. However, existing frameworks typically treat rubrics only as external evaluator disjointed from the policy's primary reasoning trace. Such design confines rubrics to post-hoc measurement, leaving them unable to actively guide the model's generation process. In this work, we introduce Think-with-Rubrics, a novel paradigm for instruction following tasks. Think-with-Rubrics integrates rubric generation into the reasoning context, transforming the rubric from an independent artifact into an internal guidance of LLM's generation. During training, LLM sequentially generates a rubric followed by a response, while a trained rubric verifier provides joint supervision by evaluating the consistency between the answer and the self-generated / golden rubrics. Experiments across multiple benchmarks demonstrate that Think-with-Rubrics consistently outperforms the Rubric-as-Reward baseline supervised by golden rubrics by an average of 3.87 points. We have also discussed the mechanism by which Think-with-Rubrics enhances model performance. Experimental results demonstrate that supervision from golden rubrics and self-generated rubrics enhances the performance of Think-with-Rubrics by improving the quality of self-generated rubrics and increasing the internal consistency of responses respectively.

评分标准大模型推理生成引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。