arXiv:2603.21728cs.AIcs.CL2026-03被引 3

用检查清单反馈训练模型,让科研想法自动迭代优化。

EvoIdeator: Evolving Scientific Ideas through Checklist-Grounded Reinforcement Learning

  • 以检查清单为基准生成多维度奖励和细粒度语言反馈
  • 在科学质量指标上超越更大规模前沿模型
  • 无需微调即可适应外部反馈,适合自主科研系统

科学创意生成是自主知识发现的核心,但将初始概念演化为高质量研究提案所需的迭代过程,对大语言模型仍具挑战。现有强化学习方法依赖基于评分量表的全局质量评分,缺乏可操作的细节。语言式优化方法则通常仅限于推理时提示,未针对模型内化批评进行专门优化。为此,我们提出EvoIdeator框架,通过将强化学习目标与检查清单接地反馈对齐,实现科学创意的演化。该框架利用结构化裁判模型生成两类协同信号:(1)用于多维优化的词典序奖励,(2)关于可验证性、可行性与方法严谨性的逐段批评。将这些信号融入强化学习循环后,策略模型可在优化与推理中系统性使用精确反馈。大量实验表明,基于Qwen3-4B构建的EvoIdeator,在关键科学指标上显著优于更大规模的前沿模型。关键的是,所学策略在不经过进一步微调的情况下,能良好泛化至多种外部反馈源,为自修正自主创意生成提供了可扩展且严谨的路径。

原文摘要 · Abstract (English)

Scientific idea generation is a cornerstone of autonomous knowledge discovery, yet the iterative evolution required to transform initial concepts into high-quality research proposals remains a formidable challenge for Large Language Models (LLMs). Existing Reinforcement Learning (RL) paradigms often rely on rubric-based scalar rewards that provide global quality scores but lack actionable granularity. Conversely, language-based refinement methods are typically confined to inference-time prompting, targeting models that are not explicitly optimized to internalize such critiques. To bridge this gap, we propose \textbf{EvoIdeator}, a framework that facilitates the evolution of scientific ideas by aligning the RL training objective with \textbf{checklist-grounded feedback}. EvoIdeator leverages a structured judge model to generate two synergistic signals: (1) \emph{lexicographic rewards} for multi-dimensional optimization, and (2) \emph{fine-grained language feedback} that offers span-level critiques regarding grounding, feasibility, and methodological rigor. By integrating these signals into the RL loop, we condition the policy to systematically utilize precise feedback during both optimization and inference. Extensive experiments demonstrate that EvoIdeator, built on Qwen3-4B, significantly outperforms much larger frontier models across key scientific metrics. Crucially, the learned policy exhibits strong generalization to diverse external feedback sources without further fine-tuning, offering a scalable and rigorous path toward self-refining autonomous ideation.

科研自动化强化学习创意生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。