arXiv:2410.15403cs.CVcs.AI2024-10

整合影像分析与科室知识库,实现医学诊断的多模态智能辅助

MMDS: A Multimodal Medical Diagnosis System Integrating Image Analysis and Knowledge-based Departmental Consultation

  • 用多模态模型分析医疗影像和面部表情,精准识别情绪与面瘫
  • 面瘫分级准确率达83.3%,比GPT-4o高30个百分点
  • 按科室路由知识库,提升大模型诊断检索精度

我们提出MMDS系统,具备识别医疗图像与患者面部特征并生成专业诊断的能力。系统包含两部分:第一部分为医学影像与视频分析,训练了专用多模态医学模型,可在FER2013情感识别数据集上达到72.59%准确率,对“高兴”情绪识别达91.1%;在面瘫识别中准确率达92%,较GPT-4o高出30%。基于此模型,开发了面瘫运动视频解析器,在30例面瘫患者视频测试中,分级准确率为83.3%。第二部分为专业诊断生成,采用融合医学知识库的大语言模型,核心创新在于构建科室专属知识库路由机制,通过模型按科室分类数据,并在检索时自动选择对应知识库,显著提升RAG流程中的检索准确性。

原文摘要 · Abstract (English)

We present MMDS, a system capable of recognizing medical images and patient facial details, and providing professional medical diagnoses. The system consists of two core components:The first component is the analysis of medical images and videos. We trained a specialized multimodal medical model capable of interpreting medical images and accurately analyzing patients' facial emotions and facial paralysis conditions. The model achieved an accuracy of 72.59% on the FER2013 facial emotion recognition dataset, with a 91.1% accuracy in recognizing the "happy" emotion. In facial paralysis recognition, the model reached an accuracy of 92%, which is 30% higher than that of GPT-4o. Based on this model, we developed a parser for analyzing facial movement videos of patients with facial paralysis, achieving precise grading of the paralysis severity. In tests on 30 videos of facial paralysis patients, the system demonstrated a grading accuracy of 83.3%.The second component is the generation of professional medical responses. We employed a large language model, integrated with a medical knowledge base, to generate professional diagnoses based on the analysis of medical images or videos. The core innovation lies in our development of a department-specific knowledge base routing management mechanism, in which the large language model categorizes data by medical departments and, during the retrieval process, determines the appropriate knowledge base to query. This significantly improves retrieval accuracy in the RAG (retrieval-augmented generation) process.

医学AI多模态面瘫识别知识路由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。