让AI医生能看图问诊,像真人一样精准分析多模态医疗信息。
Advancing Conversational Diagnostic AI with Multimodal Reasoning
- 引入动态状态感知对话框架,根据患者状态变化智能追问。
- 在105个场景中,多模态表现优于94%的临床医生,诊断准确率领先。
- 适合医疗AI研发、远程诊疗系统设计者参考。
大型语言模型(LLMs)在诊断对话中展现巨大潜力,但评估多局限于纯文本交互,与远程医疗真实需求脱节。即时通讯平台支持医患上传并讨论多模态医疗资料,但现有模型对这类数据的推理能力及对话质量尚不明确。本文通过增强艺术化医疗智能探索者(AMIE)的多模态数据收集与推理能力,实现更接近资深医生的结构化问诊流程。系统基于Gemini 2.0 Flash构建状态感知对话机制,依据中间输出动态调整对话流,关键追问由患者状态不确定性驱动。在随机盲法OSCE式测试中,105个涵盖皮肤照片、心电图、临床文档等多模态资料的案例对比显示:专家评估中,AMIE在9项多模态维度中7项优于初级医师,在32项非多模态维度中29项胜出(包括诊断准确性)。结果表明,多模态对话式诊断AI取得显著进展,但实际落地仍需进一步研究。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated great potential for conducting diagnostic conversations but evaluation has been largely limited to language-only interactions, deviating from the real-world requirements of remote care delivery. Instant messaging platforms permit clinicians and patients to upload and discuss multimodal medical artifacts seamlessly in medical consultation, but the ability of LLMs to reason over such data while preserving other attributes of competent diagnostic conversation remains unknown. Here we advance the conversational diagnosis and management performance of the Articulate Medical Intelligence Explorer (AMIE) through a new capability to gather and interpret multimodal data, and reason about this precisely during consultations. Leveraging Gemini 2.0 Flash, our system implements a state-aware dialogue framework, where conversation flow is dynamically controlled by intermediate model outputs reflecting patient states and evolving diagnoses. Follow-up questions are strategically directed by uncertainty in such patient states, leading to a more structured multimodal history-taking process that emulates experienced clinicians. We compared AMIE to primary care physicians (PCPs) in a randomized, blinded, OSCE-style study of chat-based consultations with patient actors. We constructed 105 evaluation scenarios using artifacts like smartphone skin photos, ECGs, and PDFs of clinical documents across diverse conditions and demographics. Our rubric assessed multimodal capabilities and other clinically meaningful axes like history-taking, diagnostic accuracy, management reasoning, communication, and empathy. Specialist evaluation showed AMIE to be superior to PCPs on 7/9 multimodal and 29/32 non-multimodal axes (including diagnostic accuracy). The results show clear progress in multimodal conversational diagnostic AI, but real-world translation needs further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。