用大模型多智能体模拟多学科会诊,提升慢性病共病治疗推荐安全性。
Lessons Learned from Evaluation of LLM based Multi-agents in Safer Therapy Recommendation
- 设计多智能体系统模拟医生会诊,通过对话解决用药冲突。
- 单个大模型表现已达多学科团队水平,但建议仍不完整且存冗余用药。
- 新增临床目标达成率与用药负担指标,更贴合真实医疗需求。
针对多重慢性病患者治疗推荐中的用药冲突风险,现有决策支持系统存在可扩展性瓶颈。受全科医生处理共病患者时偶发多学科团队(MDT)协作的启发,本研究探索了基于大语言模型(LLM)的多智能体系统(MAS)在安全治疗推荐中的可行性与价值。设计了单智能体与模拟MDT决策的多智能体框架,通过智能体间讨论化解医学冲突。在基准病例上评估治疗规划任务,对比了多智能体与单智能体方法及真实世界基准。研究创新在于提出超越技术精确率与召回率的评价指标,可检验建议是否达成临床目标及药物负担。结果表明,当前大模型下,单智能体全科医生表现等同于多学科团队;最优模型虽能正确覆盖所有临床目标,但建议不完整,部分模型还引入冗余药物,导致药物-疾病或药-药相互作用冲突。
原文摘要 · Abstract (English)
Therapy recommendation for chronic patients with multimorbidity is challenging due to risks of treatment conflicts. Existing decision support systems face scalability limitations. Inspired by the way in which general practitioners (GP) manage multimorbidity patients, occasionally convening multidisciplinary team (MDT) collaboration, this study investigated the feasibility and value of using a Large Language Model (LLM)-based multi-agent system (MAS) for safer therapy recommendations. We designed a single agent and a MAS framework simulating MDT decision-making by enabling discussion among LLM agents to resolve medical conflicts. The systems were evaluated on therapy planning tasks for multimorbidity patients using benchmark cases. We compared MAS performance with single-agent approaches and real-world benchmarks. An important contribution of our study is the definition of evaluation metrics that go beyond the technical precision and recall and allow the inspection of clinical goals met and medication burden of the proposed advices to a gold standard benchmark. Our results show that with current LLMs, a single agent GP performs as well as MDTs. The best-scoring models provide correct recommendations that address all clinical goals, yet the advices are incomplete. Some models also present unnecessary medications, resulting in unnecessary conflicts between medication and conditions or drug-drug interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。