首个面向中医的多模态问答基准,涵盖52000+题,支持图文视频综合评估。
TCM-Ladder: A Benchmark for Multimodal Question Answering on Traditional Chinese Medicine
- 构建覆盖中医7大领域的多模态问答数据集,含图文视频等多源信息。
- 包含超5.2万道题目,涵盖单选、多选、填空、问诊对话等多种题型。
- 提出专用评分方法Ladder-Score,精准评估术语与语义表达质量。
中医作为有效替代医学正受关注。近年来,针对中医的大语言模型快速发展,亟需客观全面的评估框架。然而现有数据集范围有限,以文本为主,缺乏统一标准化的多模态问答基准。为此,我们提出TCM-Ladder,首个专为评估中医大模型设计的综合性多模态问答数据集。涵盖中医基础理论、诊断、方剂、内科、外科、药学、儿科等核心领域。除文本外,还融合图像与视频等多模态内容。通过自动化与人工筛选结合构建,共包含超过52,000道题目,题型包括单选、多选、填空、问诊对话及视觉理解。我们在该数据集上训练推理模型,并对比九个主流通用领域与五个领先中医专用大模型的表现。此外,提出Ladder-Score评估方法,专门衡量答案在术语使用与语义表达上的质量。据我们所知,这是首个系统性地在统一多模态基准上评估主流通用与中医专用大模型的工作。数据集与排行榜已公开于https://tcmladder.com,将持续更新。
原文摘要 · Abstract (English)
Traditional Chinese Medicine (TCM), as an effective alternative medicine, has been receiving increasing attention. In recent years, the rapid development of large language models (LLMs) tailored for TCM has highlighted the urgent need for an objective and comprehensive evaluation framework to assess their performance on real-world tasks. However, existing evaluation datasets are limited in scope and primarily text-based, lacking a unified and standardized multimodal question-answering (QA) benchmark. To address this issue, we introduce TCM-Ladder, the first comprehensive multimodal QA dataset specifically designed for evaluating large TCM language models. The dataset covers multiple core disciplines of TCM, including fundamental theory, diagnostics, herbal formulas, internal medicine, surgery, pharmacognosy, and pediatrics. In addition to textual content, TCM-Ladder incorporates various modalities such as images and videos. The dataset was constructed using a combination of automated and manual filtering processes and comprises over 52,000 questions. These questions include single-choice, multiple-choice, fill-in-the-blank, diagnostic dialogue, and visual comprehension tasks. We trained a reasoning model on TCM-Ladder and conducted comparative experiments against nine state-of-the-art general domain and five leading TCM-specific LLMs to evaluate their performance on the dataset. Moreover, we propose Ladder-Score, an evaluation method specifically designed for TCM question answering that effectively assesses answer quality in terms of terminology usage and semantic expression. To the best of our knowledge, this is the first work to systematically evaluate mainstream general domain and TCM-specific LLMs on a unified multimodal benchmark. The datasets and leaderboard are publicly available at https://tcmladder.com and will be continuously updated.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。