首个动态可扩展的中医大模型评测基准,推动中医AI发展
TCM-Eval: An Expert-Level Dynamic and Extensible Benchmark for Traditional Chinese Medicine
- 基于执业医师考试题构建动态评测集,专家验证确保质量
- 自迭代推理增强生成高质量问答对,实现数据与模型协同进化
- 推出专用于中医的领先大模型ZhiMingTang,超越人类执业标准
大语言模型在现代医学中表现卓越,但在中医领域因缺乏标准化评测基准和高质量训练数据而进展缓慢。为此,我们提出TCM-Eval,首个面向中医的动态可扩展评测基准,数据源自国家级中医执业资格考试,并经中医专家验证。同时,我们构建大规模训练语料库,提出自迭代思维链增强(SI-CoTE)方法,通过拒绝采样自动扩充带验证推理链的问答对,形成数据与模型协同演化的良性循环。基于此,我们开发出专为中医设计的先进大模型ZhiMingTang(ZMT),其性能显著超过人类执业者通过门槛。为促进后续研究,我们开放公共排行榜,鼓励社区参与和持续优化。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in modern medicine, yet their application in Traditional Chinese Medicine (TCM) remains severely limited by the absence of standardized benchmarks and the scarcity of high-quality training data. To address these challenges, we introduce TCM-Eval, the first dynamic and extensible benchmark for TCM, meticulously curated from national medical licensing examinations and validated by TCM experts. Furthermore, we construct a large-scale training corpus and propose Self-Iterative Chain-of-Thought Enhancement (SI-CoTE) to autonomously enrich question-answer pairs with validated reasoning chains through rejection sampling, establishing a virtuous cycle of data and model co-evolution. Using this enriched training data, we develop ZhiMingTang (ZMT), a state-of-the-art LLM specifically designed for TCM, which significantly exceeds the passing threshold for human practitioners. To encourage future research and development, we release a public leaderboard, fostering community engagement and continuous improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。