arXiv:2502.09156cs.CL2025-02被引 4

用树状知识库+自我反思,提升大模型在中医问答中的准确率

Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs

  • 构建分层树状知识库,支持跨章节信息整合检索
  • 在中医执考题上准确率提升19.85%,召回率从27%升至38%
  • 显著改善回答的安全性、一致性与可解释性,适合医学AI研发者

大型语言模型(LLMs)可利用医学知识实现智能问答,助力辅助诊断与医学人才培养。然而,传统中医(TCM)领域缺乏高效的检索增强生成(RAG)框架。本文提出树状组织的自反思检索(TOSRR)框架,通过构建具有层级结构的知识库,在推理时实现跨章节信息融合检索。以中医执业医师考试(TCM MLE)和高校经典课程考试(CCE)题目为基准数据集,实验表明:结合GPT-4后,该框架在TCM MLE上的绝对准确率提升19.85%,在CCE数据集上召回率从27%提升至38%。人工评估显示,安全、一致、可解释性、合规性与连贯性等维度总分提升18.52分。结果证明TOSRR能有效增强大模型在中医问答任务中的表现。

原文摘要 · Abstract (English)

Objectives: Large language models (LLMs) can harness medical knowledge for intelligent question answering (Q&A), promising support for auxiliary diagnosis and medical talent cultivation. However, there is a deficiency of highly efficient retrieval-augmented generation (RAG) frameworks within the domain of Traditional Chinese Medicine (TCM). Our purpose is to observe the effect of the Tree-Organized Self-Reflective Retrieval (TOSRR) framework on LLMs in TCM Q&A tasks. Materials and Methods: We introduce the novel approach of knowledge organization, constructing a tree structure knowledge base with hierarchy. At inference time, our self-reflection framework retrieves from this knowledge base, integrating information across chapters. Questions from the TCM Medical Licensing Examination (MLE) and the college Classics Course Exam (CCE) were randomly selected as benchmark datasets. Results: By coupling with GPT-4, the framework can improve the best performance on the TCM MLE benchmark by 19.85% in absolute accuracy, and improve recall accuracy from 27% to 38% on CCE datasets. In manual evaluation, the framework improves a total of 18.52 points across dimensions of safety, consistency, explainability, compliance, and coherence. Conclusion: The TOSRR framework can effectively improve LLM's capability in Q&A tasks of TCM.

中医问答知识检索大模型RAG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。