用自动更新的医学知识图谱提升大模型问诊准确率
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
- 构建可自动更新的医学知识图谱,结合外部文献检索
- 在MEDQA上达74.1% F1,MEDMCQA上达66.34%准确率
- 无需增加算力开销,适合医疗AI系统开发者
大型语言模型(LLMs)通过利用大量临床数据和医学文献显著提升了医疗问答能力。然而,医学知识快速演进以及领域资源手动更新的高成本,影响了系统的可靠性。为此,我们提出Agentic Medical Graph-RAG(AMG-RAG),一个自动化构建并持续更新医学知识图谱的框架,集成推理与外部证据检索(如PubMed、WikiSearch)。通过动态关联新发现与复杂医学概念,AMG-RAG不仅提升准确性,还增强可解释性。在MEDQA和MEDMCQA基准测试中表现优异,分别取得74.1% F1和66.34%准确率,优于许多同类模型,甚至超过规模大10至100倍的模型。关键在于,这些提升未增加计算开销,凸显了自动化知识图谱生成与外部证据检索对提供及时可信医疗洞察的关键作用。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have significantly advanced medical question-answering by leveraging extensive clinical data and medical literature. However, the rapid evolution of medical knowledge and the labor-intensive process of manually updating domain-specific resources pose challenges to the reliability of these systems. To address this, we introduce Agentic Medical Graph-RAG (AMG-RAG), a comprehensive framework that automates the construction and continuous updating of medical knowledge graphs, integrates reasoning, and retrieves current external evidence, such as PubMed and WikiSearch. By dynamically linking new findings and complex medical concepts, AMG-RAG not only improves accuracy but also enhances interpretability in medical queries. Evaluations on the MEDQA and MEDMCQA benchmarks demonstrate the effectiveness of AMG-RAG, achieving an F1 score of 74.1 percent on MEDQA and an accuracy of 66.34 percent on MEDMCQA, outperforming both comparable models and those 10 to 100 times larger. Notably, these improvements are achieved without increasing computational overhead, highlighting the critical role of automated knowledge graph generation and external evidence retrieval in delivering up-to-date, trustworthy medical insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。