用记忆与检索增强中医舌诊智能系统,提升诊断准确性。
MMIR-TCM: Memory-Integrated Multimodal Inference and Retrieval for TCM Clinical Decision Support

- 融合视觉分割与语言模型,模拟中医专家诊断流程
- 在新数据集上表现超越GPT-4o和Gemini 2.5 Flash
- 专为中医临床设计评估指标,更贴合实际需求
中医舌诊长期面临主观性强、可重复性差的问题。多模态人工智能在证候辨识与处方生成等任务中受限于视觉舌象特征与文本推理间的语义鸿沟,以及缺乏大规模标准化数据集。为此,我们提出MMIR-TCM框架,通过整合多模态大模型(MLLM)与记忆增强分割、检索增强生成(RAG),模拟中医专家的诊断过程。该框架采用三阶段设计:无需训练的Memory-SAM模块实现鲁棒舌体提取;微调后的Qwen3-VL模型生成结构化舌诊报告;基于Qwen3的RAG组件提供有证据支持的临床决策建议。系统基于我们构建的MedTCM多模态数据集进行开发与验证,并引入针对领域特点的评估指标TDEU,综合考虑语义理解与诊断重要性。实验证明,MMIR-TCM显著优于GPT-4o与Gemini 2.5 Flash等领先模型。
原文摘要 · Abstract (English)
Traditional Chinese Medicine (TCM) diagnosis, particularly through tongue inspection, faces persistent challenges in subjectivity and reproducibility. The application of multimodal artificial intelligence to TCM clinical tasks, such as syndrome differentiation and prescription generation, is significantly hampered by the semantic gap between visual tongue features and textual reasoning, as well as the lack of large-scale, standardized datasets. To address these challenges, we introduce MMIR-TCM, a novel framework that emulates the diagnostic process of TCM experts by integrating multimodal large language model(MLLM) with memory-augmented segmentation and retrieval-augmented generation (RAG). Employing a three-stage architecture, MMIR-TCM integrates a training-free Memory-SAM module for robust tongue extraction, a fine-tuned Qwen3-VL model for structured tongue diagnosis generation, and a Qwen3-based RAG component for evidence-grounded clinical decision support generation. The framework was developed and validated using MedTCM, a new large-scale multimodal dataset that we introduce specifically for advanced TCM research. To properly evaluate our framework's clinical accuracy, which existing metrics fail to capture, we also developed TDEU, a domain-specific evaluation metric incorporating semantic understanding and diagnostic importance. Our comprehensive experiments demonstrate that MMIR-TCM significantly outperforms leading models, including GPT-4o and Gemini 2.5 Flash.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。