arXiv:2607.25489cs.CV2026-07

医学领域智能体系统突破单一预测,实现多步骤临床任务协同。

Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation

论文配图:Agentic AI in medicine: architectures, applications, evaluation, and challenges for clinical translation
图 1 · 摘自论文原文
  • 构建可调用外部工具、具备记忆与纠错能力的医学智能体架构
  • 557项研究验证了其在病历分析、影像解读等场景的有效性
  • 适合医疗AI临床落地研究者关注,需加强真实场景验证

大型语言模型与多模态基础模型正推动医学AI从单一预测迈向需规划、工具调用、记忆、迭代修正及多智能体协作的多步骤临床任务。本研究通过系统性文献综述,筛选出1649条记录,最终纳入557项符合目标导向任务执行、工具使用、外部资源交互、反馈优化或多智能体协作标准的研究。这些研究涵盖单智能体工具调用、检索增强工作流、多模态智能体及多智能体系统,应用于医学问答、影像分析、电子病历处理、药物安全与临床试验预测。当前证据仍以公开基准测试、模拟环境、回顾性数据集和小规模专家评估为主,过程可靠性、证据可追溯性、不确定性、安全性、工作流影响与外部有效性评估不一致。临床转化依赖更清晰定义、可复现评估、可审计监管、互操作设计及真实世界前瞻性验证。

原文摘要 · Abstract (English)

Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iterative correction, and coordination among specialized agents. However, the scope of agentic AI in medicine remains unsettled, and current evaluation practices are not yet aligned with the requirements of clinical use. We conducted a scoping review with systematic evidence mapping across five electronic sources, screened 1,649 exportable records, and provisionally included 557 unique studies that met predefined criteria for goal-directed task execution, tool use, interaction with external resources, feedback-based refinement, or multi-agent collaboration. The included studies describe single agents that use external tools, workflows supported by retrieval and external knowledge, multimodal agents, and multi-agent systems applied to medical question answering, image interpretation, electronic health record analysis, drug safety, and clinical trial prediction. The evidence base remains dominated by public benchmarks, simulated settings, retrospective datasets, and small-scale expert evaluation. Process reliability, evidence traceability, uncertainty, safety, workflow impact, and external validity are evaluated less consistently. Clinical translation will depend on clearer definitions, reproducible evaluation, auditable oversight, interoperable system design, and prospective validation in real-world clinical workflows.

医学AI智能体系统临床落地多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。