arXiv:2602.01779cs.AI2026-02

首个统一评估大模型中医知识与临床推理能力的基准

LingLanMiDian: Systematic Evaluation of LLMs on TCM Knowledge and Clinical Reasoning

  • 构建覆盖多任务的中医标准化评测集
  • 硬样本测试揭示模型与专家间显著差距
  • 适合中医AI研究者与大模型开发者使用

大型语言模型在医学自然语言处理领域快速进步,但中医因其独特的本体、术语和推理模式,亟需领域适配的评估体系。现有中医评测数据集覆盖不全、规模有限,且评分方式不统一或依赖生成任务,难以公平比较。本文提出LingLanMiDian(LingLan)基准,一个大规模、专家标注、多任务的评测体系,涵盖知识回忆、多跳推理、信息抽取及真实临床决策。LingLan采用一致的度量设计、容忍同义词的临床标签协议、每数据集400项硬样本子集,并将诊断与治疗推荐重构为单选决策识别。我们在14个主流开源与专有大模型上进行零样本全面评估,揭示其在中医常识理解、推理与临床支持方面的优劣;尤其在硬样本测试中,模型表现远低于人类专家。通过统一基础认知与应用推理的评测标准,LingLan为中医大模型及领域专用医疗AI研究建立了一个可量化、可扩展的基准。所有数据与代码公开于https://github.com/TCMAI-BJTU/LingLan 和 http://tcmnlp.com。

原文摘要 · Abstract (English)

Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are fragmented in coverage and scale and rely on non-unified or generation-heavy scoring that hinders fair comparison. We present the LingLanMiDian (LingLan) benchmark, a large-scale, expert-curated, multi-task suite that unifies evaluation across knowledge recall, multi-hop reasoning, information extraction, and real-world clinical decision-making. LingLan introduces a consistent metric design, a synonym-tolerant protocol for clinical labels, a per-dataset 400-item Hard subset, and a reframing of diagnosis and treatment recommendation into single-choice decision recognition. We conduct comprehensive, zero-shot evaluations on 14 leading open-source and proprietary LLMs, providing a unified perspective on their strengths and limitations in TCM commonsense knowledge understanding, reasoning, and clinical decision support; critically, the evaluation on Hard subset reveals a substantial gap between current models and human experts in TCM-specialized reasoning. By bridging fundamental knowledge and applied reasoning through standardized evaluation, LingLan establishes a unified, quantitative, and extensible foundation for advancing TCM LLMs and domain-specific medical AI research. All evaluation data and code are available at https://github.com/TCMAI-BJTU/LingLan and http://tcmnlp.com.

中医AI大模型评估临床推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。