让AI主动问诊,生成更贴近真实医生的医疗报告。
ProMRVL-CAD: Proactive Dialogue System with Multi-Round Vision-Language Interactions for Computer-Aided Diagnosis
- 用知识图谱驱动主动提问,引导诊断流程
- 在MIMIC-CXR和IU-Xray数据集上报告质量更优
- 适合医学AI对话系统研究者与开发者
大语言模型(LLMs)在视觉-语言任务中展现出卓越的理解能力,但在生成可靠医疗诊断报告方面仍处于初期阶段。当前医疗LLM多采用被动交互模式,医生仅回应患者提问,很少主动分析医学影像;部分ChatBot则仅根据预设问题回复视觉输入,缺乏互动与病史考量。为此,我们提出基于LLM的主动多轮视觉-语言交互系统ProMRVL-CAD,用于生成面向患者的疾病诊断报告。该系统通过引入知识图谱构建推荐机制,实现主动对话。具体设计两个生成器:主动问题生成器(Pro-Q Gen)生成引导诊断的问题,多视觉患者文本报告生成器(MVP-DR Gen)生成高质量诊断报告。在两个公开真实数据集MIMIC-CXR和IU-Xray上评估,模型生成报告质量更优。此外,系统在低图像质量场景下仍表现稳健。我们还构建了一个模拟医患主动问诊的合成医疗对话数据集,为训练LLM提供宝贵资源。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have demonstrated extraordinary comprehension capabilities with remarkable breakthroughs on various vision-language tasks. However, the application of LLMs in generating reliable medical diagnostic reports remains in the early stages. Currently, medical LLMs typically feature a passive interaction model where doctors respond to patient queries with little or no involvement in analyzing medical images. In contrast, some ChatBots simply respond to predefined queries based on visual inputs, lacking interactive dialogue or consideration of medical history. As such, there is a gap between LLM-generated patient-ChatBot interactions and those occurring in actual patient-doctor consultations. To bridge this gap, we develop an LLM-based dialogue system, namely proactive multi-round vision-language interactions for computer-aided diagnosis (ProMRVL-CAD), to generate patient-friendly disease diagnostic reports. The proposed ProMRVL-CAD system allows proactive dialogue to provide patients with constant and reliable medical access via an integration of knowledge graph into a recommendation system. Specifically, we devise two generators: a Proactive Question Generator (Pro-Q Gen) to generate proactive questions that guide the diagnostic procedure and a Multi-Vision Patient-Text Diagnostic Report Generator (MVP-DR Gen) to produce high-quality diagnostic reports. Evaluating two real-world publicly available datasets, MIMIC-CXR and IU-Xray, our model has better quality in generating medical reports. We further demonstrate the performance of ProMRVL achieves robust under the scenarios with low image quality. Moreover, we have created a synthetic medical dialogue dataset that simulates proactive diagnostic interactions between patients and doctors, serving as a valuable resource for training LLM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。