构建中医大模型评估新基准,揭示推理稳定性短板。
TCM-5CEval: Extended Deep Evaluation Benchmark for LLM's Comprehensive Clinical Research Competence in Traditional Chinese Medicine
- 设计五维中医能力评测体系,覆盖经典、临床等关键领域
- 15个主流大模型测试显示,顶尖模型在题序变化时性能下降超30%
- 首次通过排列测试暴露模型对位置偏见的脆弱性,适合中医AI研究者参考
大型语言模型在通用领域表现卓越,但在高度专业化且富含文化内涵的中医领域仍需严格评估。基于前期工作TCM-3CEval,本文提出TCM-5CEval,一个更细粒度、全面的评估基准,涵盖五大维度:(1)核心知识(TCM-Exam),(2)经典文献理解(TCM-LitQA),(3)临床决策(TCM-MRCD),(4)中药学(TCM-CMM),(5)非药物疗法(TCM-ClinNPT)。我们对15个主流大模型进行了系统评估,发现显著性能差异,其中deepseek_r1和gemini_2_5_pro表现领先。结果表明,模型虽能较好记忆基础医学知识,但在解读经典文本的阐释性问题上表现不佳。关键的是,排列一致性测试揭示了普遍存在的推理脆弱性:所有模型,包括最高分模型,在不同选项顺序下均出现显著性能下降,表明其推理严重依赖位置偏见,缺乏稳健理解。TCM-5CEval不仅为中医大模型能力提供更精细诊断工具,也暴露出其推理稳定性的根本缺陷。为推动研究与标准化比较,该基准已上线Medbench平台,加入‘中医综合能力深度挑战’专项赛道。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated exceptional capabilities in general domains, yet their application in highly specialized and culturally-rich fields like Traditional Chinese Medicine (TCM) requires rigorous and nuanced evaluation. Building upon prior foundational work such as TCM-3CEval, which highlighted systemic knowledge gaps and the importance of cultural-contextual alignment, we introduce TCM-5CEval, a more granular and comprehensive benchmark. TCM-5CEval is designed to assess LLMs across five critical dimensions: (1) Core Knowledge (TCM-Exam), (2) Classical Literacy (TCM-LitQA), (3) Clinical Decision-making (TCM-MRCD), (4) Chinese Materia Medica (TCM-CMM), and (5) Clinical Non-pharmacological Therapy (TCM-ClinNPT). We conducted a thorough evaluation of fifteen prominent LLMs, revealing significant performance disparities and identifying top-performing models like deepseek\_r1 and gemini\_2\_5\_pro. Our findings show that while models exhibit proficiency in recalling foundational knowledge, they struggle with the interpretative complexities of classical texts. Critically, permutation-based consistency testing reveals widespread fragilities in model inference. All evaluated models, including the highest-scoring ones, displayed a substantial performance degradation when faced with varied question option ordering, indicating a pervasive sensitivity to positional bias and a lack of robust understanding. TCM-5CEval not only provides a more detailed diagnostic tool for LLM capabilities in TCM but aldso exposes fundamental weaknesses in their reasoning stability. To promote further research and standardized comparison, TCM-5CEval has been uploaded to the Medbench platform, joining its predecessor in the "In-depth Challenge for Comprehensive TCM Abilities" special track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。