arXiv:2508.14706cs.CLcs.AI2025-08被引 7

首个专为中医设计的多模态大模型,能综合视觉、听觉等多感官信息诊断。

ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine

  • 构建首个面向中医的多模态大模型,融合文本、图像、音频与生理信号。
  • 在1.2万张图像和200小时音频上训练,超越同类模型的中医视觉理解能力。
  • 适合中医研究者、智能医疗开发者,推动中医智能化诊断发展。

尽管大语言模型在多个领域取得成功,其在中医领域的潜力仍因两大障碍未被充分挖掘:一是高质量中医数据稀缺,二是中医诊断本身具有多模态特性,涉及望、闻、问、切等多种感官输入,超出了传统语言模型的能力范围。为此,我们提出首个面向中医的多模态大模型——ShizhenGPT。为缓解数据稀缺问题,我们构建了迄今最大的中医数据集,包含超过100GB文本和200GB多模态数据,涵盖120万张图像、200小时音频及生理信号。ShizhenGPT经过预训练与指令微调,具备深度中医知识与跨模态推理能力。评估方面,我们收集近年国家中医执业资格考试题,并建立中药识别与视觉诊断基准。实验表明,ShizhenGPT在同等规模模型中表现最优,甚至可比肩更大规模的闭源模型。尤其在中医视觉理解方面,领先现有多模态大模型,展现出对声音、脉搏、气味与视觉的统一感知能力,为实现中医全模态智能诊断奠定基础。数据集、模型与代码均已公开,旨在激发该领域的进一步探索。

原文摘要 · Abstract (English)

Despite the success of large language models (LLMs) in various domains, their potential in Traditional Chinese Medicine (TCM) remains largely underexplored due to two critical barriers: (1) the scarcity of high-quality TCM data and (2) the inherently multimodal nature of TCM diagnostics, which involve looking, listening, smelling, and pulse-taking. These sensory-rich modalities are beyond the scope of conventional LLMs. To address these challenges, we present ShizhenGPT, the first multimodal LLM tailored for TCM. To overcome data scarcity, we curate the largest TCM dataset to date, comprising 100GB+ of text and 200GB+ of multimodal data, including 1.2M images, 200 hours of audio, and physiological signals. ShizhenGPT is pretrained and instruction-tuned to achieve deep TCM knowledge and multimodal reasoning. For evaluation, we collect recent national TCM qualification exams and build a visual benchmark for Medicinal Recognition and Visual Diagnosis. Experiments demonstrate that ShizhenGPT outperforms comparable-scale LLMs and competes with larger proprietary models. Moreover, it leads in TCM visual understanding among existing multimodal LLMs and demonstrates unified perception across modalities like sound, pulse, smell, and vision, paving the way toward holistic multimodal perception and diagnosis in TCM. Datasets, models, and code are publicly available. We hope this work will inspire further exploration in this field.

中医智能多模态大模型视觉诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。