arXiv:2509.19834cs.CLcs.AI2025-09

天慧模型专为中医场景打造,实现知识高效保存与应用

TianHui: A Domain-Specific Large Language Model for Diverse Traditional Chinese Medicine Scenarios

  • 融合中医语料与领域知识,采用两阶段训练优化性能
  • 在12个评测集上6项排名第一,整体表现领先
  • 适合中医研究、临床辅助与知识数字化从业者

针对中医领域专用大模型在科研中存在适应性差、评估数据不足和计算资源受限的问题,本文提出天慧(TianHui)——一个基于上下文数据融合与领域知识整合的中医专用大语言模型。构建了大规模中医语料库(0.97GB无监督数据 + 611,312组问答对),采用两阶段训练策略,结合QLoRA、DeepSpeed Stage 2与Flash Attention 2技术。在12个基准测试中,天慧在六个数据集(APQ、TCMCD、HFR、HCCA、DHPE、TLAW)上所有指标位列前三,其余六个(TCMEE、APR、GCPMI、TCMKQA、TCMRC、ADTG)取得最优结果。最优配置为LoRA秩=128,缩放系数=256,训练轮次=4,丢弃率=0.2,最大长度=2048。该模型支持中医知识系统性保存与规模化应用,所有资源已开源。

原文摘要 · Abstract (English)

Domain-specific LLMs in TCM face limitations in research settings due to constrained adaptability, insufficient evaluation datasets, and limited computational resources. This study presents TianHui, a specialized TCM LLM built through contextual data integration and domain knowledge fusion. We constructed a large-scale TCM corpus (0.97GB unsupervised data + 611,312 QA pairs) and employed a two-stage training strategy with QLoRA, DeepSpeed Stage 2, and Flash Attention 2. Evaluation on 12 benchmarks showed TianHui ranked top-three in all metrics for six datasets (APQ, TCMCD, HFR, HCCA, DHPE, TLAW) and achieved top results in the other six (TCMEE, APR, GCPMI, TCMKQA, TCMRC, ADTG). Optimal configuration was identified as LoRA rank=128, alpha=256, epoch=4, dropout=0.2, max length=2048. TianHui enables systematic preservation and scalable application of TCM knowledge. All resources are open-sourced.

中医AI大模型知识融合自然语言处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。