arXiv:2603.23520cs.CLcs.AI2026-03

用轻量大模型系统化复制名医诊疗思维,实现可标准化的临床专家能力迁移。

From Physician Expertise to Clinical Agents: Preserving, Standardizing, and Scaling Physicians' Medical Expertise with Lightweight LLM

  • 构建五阶段框架,从多位名医经验中提取诊疗哲学与个体化规则
  • 单模型在1.5B参数下完成七类任务,性能媲美超大规模模型
  • 揭示大模型评判存在细粒度偏差,需结合医生判断确保可靠性

医学是通过长期观察和临床实践不断精进的经验学科。医师通过反复应用、反思与改进形成个性化诊疗方法,但其成果差异大,名医知识积累慢且难以规模化传播,导致高质量临床能力稀缺。为此,我们提出Med-Shicheng框架,使大语言模型能系统学习并标准化传递五位国家级中医大师的诊断治疗理念与个体化调整治疗规则。基于天翼模型,该框架包含五个阶段:收集多源资料,训练单一模型融合五位名医的知识体系,覆盖病因病机分析、证候诊断、治则选择、处方生成、处方解释、症状演变与方药调整、临床建议等七项任务。在Qwen2.5-1.5B-Base上实现部署,仅需资源受限的GPU,性能接近DeepSeek-R1与GPT-5。同时评估了大模型作为评判者与医师评价的一致性:自动化评判能跟踪整体趋势,但在个体化细节上存在偏差,提示当真实标准缺失时,仍需医生参与,并需开发领域适配的评判模型。

原文摘要 · Abstract (English)

Medicine is an empirical discipline refined through long-term observation and the messy, high-variance reality of clinical practice. Physicians build diagnostic and therapeutic competence through repeated cycles of application, reflection, and improvement, forming individualized methodologies. Yet outcomes vary widely, and master physicians' knowledge systems are slow to develop and hard to transmit at scale, contributing to the scarcity of high-quality clinical expertise. To address this, we propose Med-Shicheng, a general framework that enables large language models to systematically learn and transfer distinguished physicians' diagnostic-and-therapeutic philosophy and case-dependent adaptation rules in a standardized way. Built on Tianyi, Med-Shicheng consists of five stages. We target five National Masters of Chinese Medicine or distinguished TCM physicians, curate multi-source materials, and train a single model to internalize all five knowledge systems across seven tasks, including etiology-pathogenesis analysis, syndrome diagnosis, treatment principle selection, prescription generation, prescription explanation, symptom evolution with regimen adjustment, and clinical advice. Implemented on Qwen2.5-1.5B-Base, Med-Shicheng runs on resource-constrained GPUs while achieving performance comparable to DeepSeek-R1 and GPT-5. We also examine the reliability of LLM-as-a-judge versus physician evaluation: automated judging tracks overall trends but shows bias on fine-grained individualized distinctions, highlighting the need for physician involvement when ground truth is unavailable and for domain-adapted judge models.

医学AI轻量模型专家系统中医智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。