arXiv:2506.01252cs.CLcs.AI2025-06被引 10

构建首个中医多任务评估基准,覆盖知识、推理与安全三大核心能力。

MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine

  • 联合中医专家设计12个子数据集,涵盖诊断、开方等真实临床场景。
  • 现有大模型在基础问答表现良好,但在临床推理和用药安全上严重不足。
  • 适合研究中医AI、医疗大模型评估及可信医疗系统开发的团队使用。

中医是拥有数千年临床经验的完整医学体系,在全球健康中发挥重要作用,尤其在东亚地区。然而,中医隐含的推理逻辑、多样化的文本形式以及缺乏标准化,给计算建模与评估带来重大挑战。大语言模型在通用医学领域已展现强大潜力,但其在中医领域的系统性评估仍不充分。现有基准多局限于事实问答,缺乏领域特定任务与临床真实性。为此,我们提出MTCMB——一个面向中医知识、推理与安全的多任务评估基准。该基准由认证中医专家参与共建,包含12个子数据集,覆盖五大类别:知识问答、语言理解、诊断推理、处方生成与安全评估。数据融合真实病例记录、国家级执业考试题与经典文献,提供真实且全面的测试环境。初步结果显示,当前大模型在基础知识上表现尚可,但在临床推理、处方规划与用药安全方面存在明显短板。这些发现凸显了像MTCMB这样领域对齐基准的紧迫性,有助于推动更专业、可信的医疗AI发展。所有数据集、代码与评估工具均开源,详见:https://github.com/Wayyuanyuan/MTCMB。

原文摘要 · Abstract (English)

Traditional Chinese Medicine (TCM) is a holistic medical system with millennia of accumulated clinical experience, playing a vital role in global healthcare-particularly across East Asia. However, the implicit reasoning, diverse textual forms, and lack of standardization in TCM pose major challenges for computational modeling and evaluation. Large Language Models (LLMs) have demonstrated remarkable potential in processing natural language across diverse domains, including general medicine. Yet, their systematic evaluation in the TCM domain remains underdeveloped. Existing benchmarks either focus narrowly on factual question answering or lack domain-specific tasks and clinical realism. To fill this gap, we introduce MTCMB-a Multi-Task Benchmark for Evaluating LLMs on TCM Knowledge, Reasoning, and Safety. Developed in collaboration with certified TCM experts, MTCMB comprises 12 sub-datasets spanning five major categories: knowledge QA, language understanding, diagnostic reasoning, prescription generation, and safety evaluation. The benchmark integrates real-world case records, national licensing exams, and classical texts, providing an authentic and comprehensive testbed for TCM-capable models. Preliminary results indicate that current LLMs perform well on foundational knowledge but fall short in clinical reasoning, prescription planning, and safety compliance. These findings highlight the urgent need for domain-aligned benchmarks like MTCMB to guide the development of more competent and trustworthy medical AI systems. All datasets, code, and evaluation tools are publicly available at: https://github.com/Wayyuanyuan/MTCMB.

中医AI多任务评估大模型评测医疗安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。