arXiv:2503.07041cs.CL2025-03被引 11

构建中医多维评估基准,检验大模型在经典与临床中的表现

TCM-3CEval: A Triaxial Benchmark for Assessing Responses from Large Language Models in Traditional Chinese Medicine

  • 从核心知识、古籍理解、临床决策三维度评测大模型
  • 中文语境模型在古籍解读与临床推理上表现更优
  • 适合关注中医AI落地与文化适配性的研究者使用

大型语言模型在自然语言处理和现代医学中表现优异,但在中医领域的评估仍不充分。为此,我们提出TCM-3CEval,一个涵盖核心知识掌握、经典文本理解与临床决策三个维度的中医评估基准。我们评估了多种模型,包括国际型(如GPT-4o)、中文型(如InternLM)及医疗专用型(如PLUSE)。结果显示:所有模型在经络穴位理论、各中医流派等专业子领域均存在局限,当前能力与临床需求仍有差距。具有中文语言与文化先验的模型在经典文本解析与临床推理中表现更佳。TCM-3CEval为中医领域AI评估设立了标准,通过多维度问题与真实病例,助力优化大模型在文化根基深厚的医学场景中的应用。该基准已上线Medbench中医赛道。

原文摘要 · Abstract (English)

Large language models (LLMs) excel in various NLP tasks and modern medicine, but their evaluation in traditional Chinese medicine (TCM) is underexplored. To address this, we introduce TCM3CEval, a benchmark assessing LLMs in TCM across three dimensions: core knowledge mastery, classical text understanding, and clinical decision-making. We evaluate diverse models, including international (e.g., GPT-4o), Chinese (e.g., InternLM), and medical-specific (e.g., PLUSE). Results show a performance hierarchy: all models have limitations in specialized subdomains like Meridian & Acupoint theory and Various TCM Schools, revealing gaps between current capabilities and clinical needs. Models with Chinese linguistic and cultural priors perform better in classical text interpretation and clinical reasoning. TCM-3CEval sets a standard for AI evaluation in TCM, offering insights for optimizing LLMs in culturally grounded medical domains. The benchmark is available on Medbench's TCM track, aiming to assess LLMs' TCM capabilities in basic knowledge, classic texts, and clinical decision-making through multidimensional questions and real cases.

中医AI大模型评测多维度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。