arXiv:2606.04507cs.CLcs.AI2026-06被引 2

让生成与评估模型一起进化,提升长文本研究质量

Self-Evolving Deep Research via Joint Generation and Evaluation

论文配图:Self-Evolving Deep Research via Joint Generation and Evaluation
图 1 · 摘自论文原文
  • 生成与评估模型共享参数,协同优化
  • 动态调整评估标准,避免优化停滞
  • 适合训练开放域研究型AI代理

大型语言模型在日常应用中日益普及,深度研究生成成为关键能力。与传统问答任务不同,深度研究报告生成缺乏明确的真值,导致奖励设计难以验证,限制了有效强化学习的应用。现有方法虽采用大模型作为评判者和依赖查询的评价标准,但仍依赖静态评估器,无法随求解器进步而调整标准,造成优化压力不足甚至饱和。为此,我们提出一种自演化协同进化训练框架SCORE,将评估器与求解器在共享参数的学习过程中紧密耦合。不将生成与评估视为独立模块,而是利用其内在关联,在单一共享参数模型内实现联合改进。为约束该过程,引入元约束机制,根据求解器表现动态调控评估环境,鼓励有效评估维度并充分探索评估空间。在多个深度研究基准上的大量实验表明,报告生成质量持续提升,证明协同进化生成与评估是训练开放式研究智能体的有前景方向。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning. Existing approaches mitigate this challenge with LLM-as-a-judge and query-dependent evaluation rubrics, but they still rely on static evaluators that cannot adapt their standards as the solver improves, leading to insufficient and eventually saturated optimization pressure. We address this limitation with a \textbf{s}elf-evolving \textbf{co}-evolutionary training framework for deep \textbf{re}search evaluation and generation (SCORE), which tightly couples an evaluator and a solver in a shared-parameter learning process. Rather than treating generation and evaluation as isolated modules, we leverage their intrinsic connection to enable joint improvement within a single shared-parameter model. To restrict this process, we introduce a meta-harness, which dynamically controls the evaluation environment based on solver performance, encouraging valid evaluation dimensions and sufficiently deep evaluator search. Extensive experiments on deep research benchmarks demonstrate consistent improvement in report generation quality, showing that co-evolving evaluation and generation is a promising direction for training open-ended research agents.

大模型自进化协同优化研究生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。