AI正从辅助科研转向全流程自动化,推动科学发现方式变革。
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery

- 构建AI驱动的科研工作流,整合文献、假设、实验与报告环节。
- 现有系统在可复现性、证据留存和跨领域适应上仍存短板。
- 适合关注科研自动化、AI协同研究及评估标准的研究者参考。
人工智能正在重塑科学研究,从单一任务辅助转向涵盖文献调研、假说生成、实验验证、报告撰写与修订的长周期工作流自动化。这一转变标志着从任务级AI向工作流级科研自动化的演进。然而当前系统仍分散割裂,在自主性、领域范围、执行环境、验证机制和人工监管方面差异显著,且面临证据保存、可复现性、弱方向拒绝、溯源追踪、跨领域鲁棒性以及责任闭环等挑战。本文通过AutoResearch框架,系统梳理了AI赋能科研工作流的全谱系发展,区分了以人类主导的提示驱动辅助(Vibe Research)与日益自主的AI主导系统。分析了控制权、证据、执行、验证与问责在工作流中的再分配,并围绕五个核心条件组织研究体系:文献与研究基础、假说形成与规划、实验与工具使用、反馈验证与评审、报告与知识传播。进一步整合了AI科学家系统、人机协作研究框架、评测基准、领域部署与开源基础设施。最后提出新颖性、有效性、影响力、可靠性与溯源性五大评估维度,指出AutoResearch的自主性具有领域依赖性,在结构化、可执行、快速验证场景中更具可信度,但在具身化、延迟响应、异构复杂或制度问责情境下受限。
原文摘要 · Abstract (English)
Scientific research is being reshaped by AI systems that move beyond isolated assistance toward longer-horizon workflows spanning literature grounding, hypothesis generation, experimentation, validation, reporting, and revision. This shift marks a transition from task-level AI for science to workflow-level research automation. Yet current systems remain fragmented, differing in autonomy, domain scope, execution environment, validation mechanism, and human oversight, while still struggling with evidence preservation, reproducibility, weak-direction rejection, provenance tracking, cross-domain robustness, and accountable scientific closure. This survey examines these developments through AutoResearch, defined as the developmental spectrum of AI-powered scientific workflow automation. Within it, Vibe Research denotes the human-steered region of prompt-based assistance and human-verified execution, whereas emerging AI-led systems coordinate larger portions of the discovery loop without achieving robust autonomy. We analyze how research systems redistribute control, evidence, execution, validation, and accountability across workflows and organize the field around five workflow conditions: literature and research grounding; hypothesis formation and planning; experimentation and tool use; feedback, validation, and review; and reporting and knowledge communication. We further synthesize AI scientist systems, mixed-initiative co-research frameworks, benchmarks, domain deployments, and open-source infrastructures. Finally, we propose five evaluation dimensions--novelty, validity, impact, reliability, and provenance--and show that AutoResearch autonomy is domain-conditioned, being more credible in structured, executable, and rapidly verifiable settings but limited in embodied, delayed, heterogeneous, ethical, or institutionally accountable contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。