AI自动生成166篇论文,覆盖67个细分领域,全程可审计。
FARS: A Fully Automated Research System Deployed at Scale

- 用多阶段智能体协作,从构思到写作全流程自动化
- 公开部署产出166篇完整论文,282份评审显示质量达标
- 适合关注AI研究自动化与可信性评估的研究者
近期自动化研究系统表明语言模型代理可生成假说、开展实验并撰写完整论文,但多数证据来自精选案例、人为设定主题或少数预定义任务。我们提出FARS(Fully Automated Research System),一个可在大规模范围内跨研究主题运行的全自动化AI for AI研究系统。FARS通过共享工作区协调各阶段智能体,自主完成选题、规划、实验与写作,记录提案、代码、日志、结果与论文等全过程产物。首次公开部署中,FARS生成了166篇完整研究论文,涵盖67个细粒度的AI/ML主题,并保留中间成果作为可审计语料库而非筛选的成功案例。我们通过282份志愿者评审对其中140篇论文进行结构化评估,包括整体评分、子项评分、完整性检查及LLM使用披露。评审结果显示,FARS在大规模公开部署中能生成具有评审价值甚至优秀的AI/ML研究成果,同时暴露了实验范围过窄、方法局限和诚信问题等常见缺陷。
原文摘要 · Abstract (English)
Recent automated research systems show that language-model agents can generate hypotheses, run experiments, and write complete manuscripts, but most evidence still comes from selected examples, human-framed topics, or a few pre-defined research tasks. We present FARS (Fully Automated Research System), a fully automated AI-for-AI research system designed to operate across research topics at scale. FARS autonomously generates and advances projects through ideation, planning, experimentation, and writing, using stage-specific agents coordinated through a shared workspace that records proposals, code, logs, results, and manuscripts. In its first public deployment, FARS produced 166 complete research papers spanning 67 fine-grained AI/ML topics while preserving intermediate artifacts as an auditable corpus rather than a curated set of successes. We evaluate this corpus with 282 structured reviews from volunteer reviewers covering 140 papers, including overall ratings, sub-scores, integrity checks, and LLM-use disclosure. The reviews indicate that FARS can produce review-worthy and occasionally strong AI/ML research artifacts in a large-scale public deployment, while also exposing recurring failure modes in narrow experimental scope, methodological limitations, and integrity issues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。