用指令微调打造中医多模态助手,提升诊断与辨识准确率。
BenCao: An Instruction-Tuned Large Language Model for Traditional Chinese Medicine
- 基于自然语言指令微调,融合经典文献与专家反馈。
- 在诊断、草药识别等任务上超越通用及专用模型表现。
- 支持舌象图像分析与多模态检索,适合中医临床与研究者使用。
中医已有两千多年历史,在全球医疗中发挥重要作用。然而,将大语言模型应用于中医仍面临整体性推理、隐含逻辑和多模态诊断线索依赖的挑战。现有中医领域大模型虽在文本理解上取得进展,但缺乏多模态整合、可解释性与临床适用性。为此,我们开发了BenCao——一个基于ChatGPT的中医多模态助手,整合结构化知识库、诊断数据与专家反馈优化。该系统通过自然语言指令微调训练,不进行参数重训,符合中医专家级推理与伦理规范。其包含超过1,000部古今中医典籍的知识库,采用场景化指令框架支持多样化交互,具备链式思维模拟机制以实现可解释推理,并通过持证中医师参与的反馈优化流程持续改进。系统连接外部API实现舌象图像分类与多模态数据库检索,动态获取诊断资源。在单选题基准测试与多模态分类任务中,BenCao在诊断、中药识别、体质辨识等任务上均优于通用及中医专用模型。截至2025年10月,该模型已部署于OpenAI GPTs Store,全球近1,000名用户使用。本研究证明,通过自然语言指令微调与多模态集成,可有效构建中医领域大模型,为生成式AI与传统医学推理对齐提供可行框架,并具备现实落地潜力。
原文摘要 · Abstract (English)
Traditional Chinese Medicine (TCM), with a history spanning over two millennia, plays a role in global healthcare. However, applying large language models (LLMs) to TCM remains challenging due to its reliance on holistic reasoning, implicit logic, and multimodal diagnostic cues. Existing TCM-domain LLMs have made progress in text-based understanding but lack multimodal integration, interpretability, and clinical applicability. To address these limitations, we developed BenCao, a ChatGPT-based multimodal assistant for TCM, integrating structured knowledge bases, diagnostic data, and expert feedback refinement. BenCao was trained through natural language instruction tuning rather than parameter retraining, aligning with expert-level reasoning and ethical norms specific to TCM. The system incorporates a comprehensive knowledge base of over 1,000 classical and modern texts, a scenario-based instruction framework for diverse interactions, a chain-of-thought simulation mechanism for interpretable reasoning, and a feedback refinement process involving licensed TCM practitioners. BenCao connects to external APIs for tongue-image classification and multimodal database retrieval, enabling dynamic access to diagnostic resources. In evaluations across single-choice question benchmarks and multimodal classification tasks, BenCao achieved superior accuracy to general-domain and TCM-domain models, particularly in diagnostics, herb recognition, and constitution classification. The model was deployed as an interactive application on the OpenAI GPTs Store, accessed by nearly 1,000 users globally as of October 2025. This study demonstrates the feasibility of developing a TCM-domain LLM through natural language-based instruction tuning and multimodal integration, offering a practical framework for aligning generative AI with traditional medical reasoning and a scalable pathway for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。