arXiv:2604.15456cs.AI2026-04

用可追溯的AI代理系统提升医学研究可信度

DeepER-Med: Advancing Deep Evidence-Based Research in Medicine Through Agentic AI

  • 构建三模块可追踪的医学研究流程,实现证据生成透明化
  • 在100个真实临床问题上优于主流平台,7例结论与临床建议一致
  • 适合医学研究者、临床决策支持场景使用

可信性与透明性是人工智能在医疗和生物医学研究中被临床采纳的关键。现有深度研究系统虽整合了多跳信息检索、推理与综合,但普遍缺乏显式的证据评估标准,易导致错误累积,难以让研究人员和临床医生判断输出可靠性。同时,当前评测方法很少针对复杂真实的医学问题进行评估。本文提出DeepER-Med——一个面向医学的深度循证研究框架,采用智能体系统,将医学研究建模为三个可追溯模块:研究规划、智能体协作与证据综合。为支持真实评估,我们构建了DeepER-MedQA数据集,包含由11位跨学科专家从真实医学研究场景中提炼的100个专家级研究问题。人工评估显示,DeepER-Med在生成新科学洞见等方面持续优于主流生产平台。通过8个真实临床案例验证,临床医生评估表明其结论在7例中与临床推荐一致,展现出在医学研究与决策支持中的实际潜力。

原文摘要 · Abstract (English)

Trustworthiness and transparency are essential for the clinical adoption of artificial intelligence (AI) in healthcare and biomedical research. Recent deep research systems aim to accelerate evidence-grounded scientific discovery by integrating AI agents with multi-hop information retrieval, reasoning, and synthesis. However, most existing systems lack explicit and inspectable criteria for evidence appraisal, creating a risk of compounding errors and making it difficult for researchers and clinicians to assess the reliability of their outputs. In parallel, current benchmarking approaches rarely evaluate performance on complex, real-world medical questions. Here, we introduce DeepER-Med, a Deep Evidence-based Research framework for Medicine with an agentic AI system. DeepER-Med frames deep medical research as an explicit and inspectable workflow of evidence-based generation, consisting of three modules: research planning, agentic collaboration, and evidence synthesis. To support realistic evaluation, we also present DeepER-MedQA, an evidence-grounded dataset comprising 100 expert-level research questions derived from authentic medical research scenarios and curated by a multidisciplinary panel of 11 biomedical experts. Expert manual evaluation demonstrates that DeepER-Med consistently outperforms widely used production-grade platforms across multiple criteria, including the generation of novel scientific insights. We further demonstrate the practical utility of DeepER-Med through eight real-world clinical cases. Human clinician assessment indicates that DeepER-Med's conclusions align with clinical recommendations in seven cases, highlighting its potential for medical research and decision support.

医学AI智能体系统循证研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。