arXiv:2603.28589cs.AIcs.LG2026-03被引 4

首个面向临床医学的自主科研框架,能生成可验证的研究想法与高质量论文。

Towards a Medical AI Scientist

  • 通过医工协同推理,将文献转化为可执行的临床研究思路。
  • 在171个案例中生成想法质量显著优于商用大模型,实验成功率更高。
  • 适合医疗AI研究者使用,推动临床科研自动化发展。

能够自动生成科学假说、开展实验并撰写论文的自主系统正成为加速发现的新范式。然而,现有AI科学家多为通用领域设计,难以适配临床医学——后者要求研究基于医学证据并处理专业数据模态。本文提出首个面向临床自主研究的Medical AI Scientist框架,通过医工协同推理机制,将大量文献综述转化为可操作的临床证据,提升研究思路的可追溯性;同时依据结构化医学写作规范与伦理政策,实现证据驱动的论文撰写。该框架支持三种研究模式:基于论文的复现、文献启发的创新和任务驱动的探索,分别对应不同层级的自动化程度。大规模评估(包括大语言模型与人类专家)表明,在171个案例、19项临床任务、6种数据模态下,其生成的研究想法质量显著优于商业LLM;方法与实现高度一致,可执行实验成功率显著提升。双盲评估显示,生成论文达到MICCAI水平,持续超越ISBI和BIBM投稿质量。本工作展示了利用AI实现医疗领域自主科学发现的巨大潜力。

原文摘要 · Abstract (English)

Autonomous systems that generate scientific hypotheses, conduct experiments, and draft manuscripts have recently emerged as a promising paradigm for accelerating discovery. However, existing AI Scientists remain largely domain-agnostic, limiting their applicability to clinical medicine, where research is required to be grounded in medical evidence with specialized data modalities. In this work, we introduce Medical AI Scientist, the first autonomous research framework tailored to clinical autonomous research. It enables clinically grounded ideation by transforming extensively surveyed literature into actionable evidence through clinician-engineer co-reasoning mechanism, which improves the traceability of generated research ideas. It further facilitates evidence-grounded manuscript drafting guided by structured medical compositional conventions and ethical policies. The framework operates under 3 research modes, namely paper-based reproduction, literature-inspired innovation, and task-driven exploration, each corresponding to a distinct level of automated scientific inquiry with progressively increasing autonomy. Comprehensive evaluations by both large language models and human experts demonstrate that the ideas generated by the Medical AI Scientist are of substantially higher quality than those produced by commercial LLMs across 171 cases, 19 clinical tasks, and 6 data modalities. Meanwhile, our system achieves strong alignment between the proposed method and its implementation, while also demonstrating significantly higher success rates in executable experiments. Double-blind evaluations by human experts and the Stanford Agentic Reviewer suggest that the generated manuscripts approach MICCAI-level quality, while consistently surpassing those from ISBI and BIBM. The proposed Medical AI Scientist highlights the potential of leveraging AI for autonomous scientific discovery in healthcare.

医疗AI自主科研论文生成临床研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。