LLM识别中药成分时依赖名字而非药理知识,准确率仅50%。
Do "New Snow Tablets" Contain Snow? Large Language Models Over-Rely on Names to Identify Ingredients of Chinese Drugs
- 用检索增强生成法,结合成分名提升识别准确性
- 在220种中药配方上,准确率从50%提升至82%
- 适合医疗AI开发者和中药数字化研究者
中医在医疗中应用日益广泛,专用大语言模型(LLMs)也应运而生以支持临床。准确识别中药成分是其基本要求。本文系统评估通用与中医专用LLMs在中药成分识别任务中的表现。结果发现,模型普遍存在误判:过度依赖药名字面意义、随意使用常见药材、面对陌生方剂行为不稳定;且无法理解验证任务。表明当前模型主要依赖名称而非系统药理知识。为此,我们提出一种聚焦成分名的检索增强生成(RAG)方法。在220个中药配方上的实验显示,该方法将成分验证准确率从约50%显著提升至82%。本工作揭示了现有中医专用LLMs的关键缺陷,并提供了提升其临床可靠性的实用方案。
原文摘要 · Abstract (English)
Traditional Chinese Medicine (TCM) has seen increasing adoption in healthcare, with specialized Large Language Models (LLMs) emerging to support clinical applications. A fundamental requirement for these models is accurate identification of TCM drug ingredients. In this paper, we evaluate how general and TCM-specialized LLMs perform when identifying ingredients of Chinese drugs. Our systematic analysis reveals consistent failure patterns: models often interpret drug names literally, overuse common herbs regardless of relevance, and exhibit erratic behaviors when faced with unfamiliar formulations. LLMs also fail to understand the verification task. These findings demonstrate that current LLMs rely primarily on drug names rather than possessing systematic pharmacological knowledge. To address these limitations, we propose a Retrieval Augmented Generation (RAG) approach focused on ingredient names. Experiments across 220 TCM formulations show our method significantly improves accuracy from approximately 50% to 82% in ingredient verification tasks. Our work highlights critical weaknesses in current TCM-specific LLMs and offers a practical solution for enhancing their clinical reliability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。