构建七维框架评估医疗大模型智能体,揭示能力分布不均。
Agentic AI in Healthcare & Medicine: A Seven-Dimensional Taxonomy for Empirical Evaluation of LLM-based Agents
- 提出包含29个子维度的七维分类体系,系统评估医疗LLM智能体。
- 76%研究实现外部知识整合,但事件触发激活仅2%完全实现。
- 多智能体架构主流,治疗规划等实操任务仍存在明显短板。
基于大语言模型(LLM)的智能体在医疗健康领域正迅速发展,涵盖电子病历分析、鉴别诊断、治疗方案制定及科研流程等多个任务。然而现有文献多为宽泛综述或单一能力深入探讨,缺乏统一评估框架。本文通过回顾49项研究,构建七维分类体系:认知能力、知识管理、交互模式、适应与学习、安全与伦理、框架类型、核心任务与子任务,共29个可操作子维度。采用明确的纳入排除标准和标注规则(完全实现、部分实现、未实现),对每项研究进行映射并报告能力分布与共现模式的量化结果。分析显示显著不对称性:知识管理中的外部知识整合约76%完全实现,而交互模式中的事件触发激活仅约8%完全实现;适应与学习中的漂移检测与缓解更是近乎缺失(约98%未实现)。架构上,多智能体设计占主导(约82%完全实现),但编排层仍以部分实现为主。核心任务中,信息类能力如医学问答与决策支持领先,而以行动和发现为导向的治疗规划与处方仍存在较大差距(约59%未实现)。
原文摘要 · Abstract (English)
Large Language Model (LLM)-based agents that plan, use tools and act has begun to shape healthcare and medicine. Reported studies demonstrate competence on various tasks ranging from EHR analysis and differential diagnosis to treatment planning and research workflows. Yet the literature largely consists of overviews which are either broad surveys or narrow dives into a single capability (e.g., memory, planning, reasoning), leaving healthcare work without a common frame. We address this by reviewing 49 studies using a seven-dimensional taxonomy: Cognitive Capabilities, Knowledge Management, Interaction Patterns, Adaptation & Learning, Safety & Ethics, Framework Typology and Core Tasks & Subtasks with 29 operational sub-dimensions. Using explicit inclusion and exclusion criteria and a labeling rubric (Fully Implemented, Partially Implemented, Not Implemented), we map each study to the taxonomy and report quantitative summaries of capability prevalence and co-occurrence patterns. Our empirical analysis surfaces clear asymmetries. For instance, the External Knowledge Integration sub-dimension under Knowledge Management is commonly realized (~76% Fully Implemented) whereas Event-Triggered Activation sub-dimenison under Interaction Patterns is largely absent (~92% Not Implemented) and Drift Detection & Mitigation sub-dimension under Adaptation & Learning is rare (~98% Not Implemented). Architecturally, Multi-Agent Design sub-dimension under Framework Typology is the dominant pattern (~82% Fully Implemented) while orchestration layers remain mostly partial. Across Core Tasks & Subtasks, information centric capabilities lead e.g., Medical Question Answering & Decision Support and Benchmarking & Simulation, while action and discovery oriented areas such as Treatment Planning & Prescription still show substantial gaps (~59% Not Implemented).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。