AI医生能像真人一样管理慢性病,且用药更精准。
Towards Conversational AI for Disease Management
- 用大模型构建会思考的诊疗助手,跟踪病情变化和治疗反应。
- 在英国指南病例中表现不输医生,用药和检查建议更符合规范。
- 擅长复杂用药题,尤其在查不到知识时仍能答对高难度问题。
尽管大型语言模型在诊断对话中展现出潜力,但其在疾病管理推理(包括疾病进展、治疗反应及安全用药处方)方面的能力仍待探索。本文通过改进此前展示的诊断能力,提出基于LLM的智能体系统AMIE,优化临床管理与对话能力,支持对疾病演化、多次就诊记录、治疗反应及药物处方的专业性推理。为确保推理基于权威临床知识,AMIE利用Gemini的长上下文能力,结合上下文检索与结构化推理,使其输出与最新临床实践指南及药品目录保持一致。在一项随机盲法虚拟客观结构化临床考试(OSCE)研究中,AMIE与21名全科医生在100个符合英国NICE指南和BMJ最佳实践指南的多阶段病例中对比,评估结果显示:由专科医生判断,AMIE在管理推理上不逊于医生,在治疗和检查建议的精确性,以及与临床指南的一致性方面表现更优。为基准化药物推理能力,我们开发了RxQA,一个源自美英两国药典的多选题基准,经认证药剂师验证。尽管医生和AMIE均能通过访问外部药物信息获益,但在更高难度题目上,AMIE优于医生。虽尚需进一步研究方可投入实际应用,但其在多项评估中的优异表现标志着对话式AI在疾病管理领域的重要进展。
原文摘要 · Abstract (English)
While large language models (LLMs) have shown promise in diagnostic dialogue, their capabilities for effective management reasoning - including disease progression, therapeutic response, and safe medication prescription - remain under-explored. We advance the previously demonstrated diagnostic capabilities of the Articulate Medical Intelligence Explorer (AMIE) through a new LLM-based agentic system optimised for clinical management and dialogue, incorporating reasoning over the evolution of disease and multiple patient visit encounters, response to therapy, and professional competence in medication prescription. To ground its reasoning in authoritative clinical knowledge, AMIE leverages Gemini's long-context capabilities, combining in-context retrieval with structured reasoning to align its output with relevant and up-to-date clinical practice guidelines and drug formularies. In a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) study, AMIE was compared to 21 primary care physicians (PCPs) across 100 multi-visit case scenarios designed to reflect UK NICE Guidance and BMJ Best Practice guidelines. AMIE was non-inferior to PCPs in management reasoning as assessed by specialist physicians and scored better in both preciseness of treatments and investigations, and in its alignment with and grounding of management plans in clinical guidelines. To benchmark medication reasoning, we developed RxQA, a multiple-choice question benchmark derived from two national drug formularies (US, UK) and validated by board-certified pharmacists. While AMIE and PCPs both benefited from the ability to access external drug information, AMIE outperformed PCPs on higher difficulty questions. While further research would be needed before real-world translation, AMIE's strong performance across evaluations marks a significant step towards conversational AI as a tool in disease management.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。