arXiv:2511.04583cs.AIcs.CL2025-11被引 8

AI模仿新手研究员,从论文出发自主提出新方法并写出新论文。

Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper

  • 基于导师论文,自动分析不足、设计新假设并迭代实验。
  • 生成的论文在评审中得分高于现有全自动系统。
  • 适合关注AI科研自动化风险与边界的研究者。

理解当前AI科学家系统(自研)的能力与风险,对保障可信且可持续的AI驱动科学进步至关重要。为此,我们开发了Jr. AI Scientist——一个最先进的自主科研系统,其模仿新手研究者的典型工作流程:给定人类导师的基准论文后,它分析其局限性,提出改进的新假设,通过迭代实验直至取得提升,并撰写包含结果的新论文。不同于以往假定全自动化或仅处理小规模代码的方法,Jr. AI Scientist遵循明确的研究流程,并利用现代编码智能体处理复杂多文件实现,产出具有科学价值的成果。实验表明,该系统成功基于真实NeurIPS、IJCV和ICLR论文,提出并实现了新方法,生成可发表的论文。评估采用AI评审员、作者自评及提交至Agents4Science(专注AI贡献的会议)进行。结果显示,由DeepReviewer评分,其论文得分高于现有全自动系统。然而,作者评估与Agents4Science反馈揭示了重要局限性,表明直接应用当前系统存在潜在风险,也指明未来研究的关键挑战。最后,我们全面报告了开发过程中识别出的各种风险。本研究厘清了当前AI科学家系统的角色与限制,为仍需人类介入的领域及系统演化可能带来的风险提供了洞见。

原文摘要 · Abstract (English)

Understanding the current capabilities and risks of AI Scientist systems (autoresearch) is essential for ensuring trustworthy and sustainable AI-driven scientific progress while preserving the integrity of the academic ecosystem. To this end, we develop Jr. AI Scientist, a state-of-the-art autonomous AI scientist system that mimics the core research workflow of a novice student researcher: Given the baseline paper from the human mentor, it analyzes its limitations, formulates novel hypotheses for improvement, iteratively experiments until improvements are achieved, and writes a paper with the results. Unlike previous approaches that assume full automation or operate on small-scale code, Jr. AI Scientist follows a well-defined research workflow and leverages modern coding agents to handle complex, multi-file implementations, leading to scientifically valuable contributions. Through our experiments, the Jr. AI Scientist successfully generated new research papers that build upon real NeurIPS, IJCV, and ICLR works by proposing and implementing novel methods. For evaluation, we conducted automated assessments using AI Reviewers, author-led evaluations, and submissions to Agents4Science, a venue dedicated to AI-driven contributions. The findings demonstrate that Jr. AI Scientist generates papers receiving higher review scores by DeepReviewer than existing fully automated systems. Nevertheless, we identify important limitations from the author evaluation and the Agents4Science reviews, indicating the potential risks of directly applying current AI Scientist systems and key challenges for future research. Finally, we comprehensively report various risks identified during development. We believe this study clarifies the current role and limitations of AI Scientist systems, offering insights into the areas that still require human expertise and the risks that may emerge as these systems evolve.

AI科研自主探索风险报告

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。