arXiv:2512.02816cs.CL2025-12被引 2

首个面向中医辨证论治的综合评估基准,填补治疗决策评价空白。

A benchmark dataset for evaluating Syndrome Differentiation and Treatment in large language models

  • 由中医专家主导构建临床案例基准,覆盖辨证与治疗全流程。
  • 在15个主流大模型上验证,首次量化处方与证候匹配度。
  • 适合医疗AI研究者、中医药智能化开发者使用。

大型语言模型在中医药领域的兴起迫切需要评估其临床应用能力。然而,中医‘辨证论治’的个体化、整体性和多样性特征使现有评估受限,多数基准仅关注知识问答或辨证准确率,忽视治疗决策评估。本文提出一个由中医专家主导的综合性临床案例基准TCM-BEST4SDT,包含四大任务:中医基础知识、医学伦理、LLM内容安全与辨证论治。通过严格的数据标注流程,采用选择题评估、裁判模型评估与专用奖励模型评估三重机制,实现对处方-证候匹配度的量化。在15个主流大模型(涵盖通用与中医领域)上的实验验证了该基准的有效性。为推动智能中医药研究,TCM-BEST4SDT现已公开。

原文摘要 · Abstract (English)

The emergence of Large Language Models (LLMs) within the Traditional Chinese Medicine (TCM) domain presents an urgent need to assess their clinical application capabilities. However, such evaluations are challenged by the individualized, holistic, and diverse nature of TCM's "Syndrome Differentiation and Treatment" (SDT). Existing benchmarks are confined to knowledge-based question-answering or the accuracy of syndrome differentiation, often neglecting assessment of treatment decision-making. Here, we propose a comprehensive, clinical case-based benchmark spearheaded by TCM experts, and a specialized reward model employed to quantify prescription-syndrome congruence. Data annotation follows a rigorous pipeline. This benchmark, designated TCM-BEST4SDT, encompasses four tasks, including TCM Basic Knowledge, Medical Ethics, LLM Content Safety, and SDT. The evaluation framework integrates three mechanisms, namely selected-response evaluation, judge model evaluation, and reward model evaluation. The effectiveness of TCM-BEST4SDT was corroborated through experiments on 15 mainstream LLMs, spanning both general and TCM domains. To foster the development of intelligent TCM research, TCM-BEST4SDT is now publicly available.

中医AI评估基准辨证论治大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。