先评估再改进,让科研智能体自动生成可执行的评分标准。
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents

- 先生成任务专属可执行评分标准,再指导研究执行
- 在三个大模型上平均提升2.08至2.95分,效果稳定
- 适合需要严谨验证的自主科研场景,如论文生成与发现
自主科学研究代理正被广泛应用于从文献综述、数据分析到实验与报告生成的全流程。然而,开放性研究任务常缺乏明确的分析方法与成功标准,导致代理可能遗漏关键分析、使用不当方法或得出证据不足的结论。为此,我们提出 AutoSciRub——一种评估优先的框架,在研究执行前生成特定任务的可执行评分标准,并用于指导执行、逐项验证及迭代修改。该框架将模糊指令分解为原子化科学目标,结合相关文献与任务可见数据,合成具体、可操作、可验证的标准。生成的评分标准使隐含的实验与证据要求显性化,为分析提供指引。修订阶段,基于评分标准的验证能识别未达标项,实现对研究报告及其支撑材料的精准优化。在 ResearchClawBench 上,AutoSciRub 在三种主干大模型下平均提升 2.08 分(固定 Codex 框架),在三种代理框架下平均提升 2.95 分(固定 DeepSeek-V4-Flash)。在 AstaBench E2E Discovery 的随机 20 任务子集上,平均提升达 16.8 分,同时保持或增加成功完成的任务数。结果表明,评估优先的引导机制为自主科学研究提供了有效且通用的控制范式。
原文摘要 · Abstract (English)
Autonomous scientific research agents are increasingly applied to end-to-end scientific workflows, including literature review, data analysis, experimentation, and report generation. However, open-ended research tasks often do not clearly specify the analyses, methods, and success criteria required to complete the task. As a result, agents may miss important analyses, use inappropriate methods, or draw conclusions that are insufficiently supported by evidence. To address the problem, we present AutoSciRub, an evaluation-first framework that induces a task-specific executable rubric before research execution, and uses it to guide execution, criterion-level verification as well as iterative revision. AutoSciRub decomposes an underspecified instruction into atomic scientific goals, grounds them in relevant literature and task-visible data, and synthesizes specific, actionable, and verifiable criteria. The resulting rubric makes implicit experimental and evidential requirements explicit, providing guidance for experiments and analyses. During revision, rubric-guided verification identifies unmet criteria and enables targeted refinement of the research report and its supporting artifacts. On ResearchClawBench, AutoSciRub consistently improves all tested configurations, with an average gain of 2.08 points across three backbone LLMs under the fixed Codex harness and 2.95 points across three agent harnesses using a fixed DeepSeek-V4-Flash backbone. On a randomly sampled 20-task subset of AstaBench E2E Discovery, AutoSciRub further achieves an average improvement of 16.8 points across three agent harnesses, while maintaining or increasing the number of successfully completed tasks. These results demonstrate that evaluation-first guidance provides an effective and generalizable control mechanism for autonomous scientific research (Code: https://github.com/zjunlp/AutoSciRub).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。