用论文自动生成评分标准,训练AI制定更符合要求的研究计划。
Training AI Co-Scientists Using Rubric Rewards
- 用论文提取研究目标和评分标准,构建自动训练数据集
- 通过自评机制在无真人监督下提升计划生成质量
- 在机器学习、医学等领域均有效,适合想高效规划研究的人
AI协作者正成为辅助研究人员实现研究目标的工具,其核心能力是根据研究目标与约束生成研究计划。当前语言模型难以生成完全符合所有约束和隐含要求的计划。本文提出一种新方法:利用海量已有论文构建可扩展、多样化的训练语料库,自动提取研究目标及对应评分标准。通过强化学习进行微调,使用初始模型的冻结版本作为自评器,评分标准制造生成-验证差距,实现无需人工干预的性能提升。在225小时的人类专家评估中,70%的研究目标下,专家更偏好经过微调的Qwen3-30B-A3B模型生成的计划;84%的自动生成评分标准被认可。该方法还扩展至医学论文与arXiv预印本,由前沿模型组成的评审团评估显示,相对提升12%-22%,具备显著跨领域泛化能力,即使在缺乏执行反馈的医疗研究场景中仍有效。结果证明,这种自动化训练范式为提升通用型AI协作者提供了可行路径。
原文摘要 · Abstract (English)
AI co-scientists are emerging as a tool to assist human researchers in achieving their research goals. A crucial feature of these AI co-scientists is the ability to generate a research plan given a set of aims and constraints. The plan may be used by researchers for brainstorming, or may even be implemented after further refinement. However, language models currently struggle to generate research plans that follow all constraints and implicit requirements. In this work, we study how to leverage the vast corpus of existing research papers to train language models that generate better research plans. We build a scalable, diverse training corpus by automatically extracting research goals and goal-specific grading rubrics from papers across several domains. We then train models for research plan generation via reinforcement learning with self-grading. A frozen copy of the initial policy acts as the grader during training, with the rubrics creating a generator-verifier gap that enables improvements without external human supervision. To validate this approach, we conduct a study with human experts for machine learning research goals, spanning 225 hours. The experts prefer plans generated by our finetuned Qwen3-30B-A3B model over the initial model for 70% of research goals, and approve 84% of the automatically extracted goal-specific grading rubrics. To assess generality, we also extend our approach to research goals from medical papers, and new arXiv preprints, evaluating with a jury of frontier models. Our finetuning yields 12-22% relative improvements and significant cross-domain generalization, proving effective even in problem settings like medical research where execution feedback is infeasible. Together, these findings demonstrate the potential of a scalable, automated training recipe as a step towards improving general AI co-scientists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。