arXiv:2602.03619cs.CL2026-02被引 15

用人类偏好训练可定制的报告评分标准,提升长文本生成质量。

Learning Query-Specific Rubrics from Human Preferences for DeepResearch Report Generation

  • 基于人类偏好数据,用强化学习训练查询专属评分生成器。
  • 在测试集上优于通用或人工设计的评分标准,能更好区分优劣报告。
  • 适用于需要高质量长文生成的研究型AI系统,尤其适合多智能体框架。

当前构建可靠的DeepResearch风格长篇报告仍具挑战性,因训练与评估缺乏可验证的奖励信号。因此,基于评分标准的评估成为常见做法。然而,现有方法要么依赖粗糙的预设评分标准,缺乏细粒度;要么依赖人工构建的查询专属评分标准,成本高且难以扩展。本文提出一种流水线方法,训练面向DeepResearch报告生成的偏好驱动型查询专属评分生成器。我们首先构建了一个标注了人类对成对报告偏好的DeepResearch查询数据集,并通过融合偏好一致性、格式有效性及基于大模型的评分评估的混合奖励信号,使用强化学习训练评分生成器。评估分两阶段进行:首先,在保留的人类偏好测试集上,所学评分标准比通用提示、微调训练或人工构造的评分标准更有效地区分优选与拒选报告;其次,当作为奖励信号用于训练DeepResearch系统时,该评分生成器在DeepResearch Bench上的单智能体ReAct框架和复杂多智能体工作流中均带来显著性能提升。

原文摘要 · Abstract (English)

Nowadays, developing reliable DeepResearch-style long-form report generation remains challenging, as training and evaluation lack verifiable reward signals. Accordingly, rubric-based evaluation has become a common practice. However, existing approaches either rely on coarse, pre-defined rubrics that lack sufficient granularity or depend on manually constructed query-specific rubrics that are costly and difficult to scale. In this paper, we propose a pipeline to train preference-grounded query-specific rubric generators tailored for DeepResearch report generation. We first construct a dataset of DeepResearch-style queries annotated with human preferences over paired reports, and train rubric generators via reinforcement learning with a hybrid reward combining preference consistency, format validity, and LLM-based rubric evaluation. We evaluate the resulting rubric generators in two stages. First, on a held-out human-preference test set, the learned rubrics discriminate preferred from rejected reports more effectively than generic, prompted, or SFT-trained rubric alternatives. Second, when used as reward signals to train DeepResearch systems, our rubric generators yield substantial performance gains under both a simple single-agent ReAct framework and a complex multi-agent workflow on the DeepResearch Bench.

报告生成偏好学习评分标准强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。