用外部知识增强大模型,提升胰腺癌分期准确率。
Enhancing Pancreatic Cancer Staging with Large Language Models: The Role of Retrieval-Augmented Generation
- 引入外部医学指南作为知识源,通过检索增强生成提升推理能力。
- 使用真实病例测试,检索增强模型分期准确率达70%,显著高于无检索模型的35%。
- 可展示引用来源,帮助医生理解判断依据,适合临床辅助诊断场景。
目的:检索增强生成(RAG)通过从可靠外部知识(REK)中检索相关信息,提升大语言模型(LLM)的功能性和可靠性。尽管RAG在放射学领域受到关注,我们此前已报道使用带RAG的NotebookLM在肺癌分期中的有效性。但由于对比的LLM与其内部模型不同,难以确定优势是来自RAG本身还是模型差异。为更清晰地评估RAG的影响并验证其在多种癌症中的适用性,本研究在胰腺癌分期任务中,将NotebookLM与自身内部模型Gemini 2.0 Flash进行对比。方法:以日本胰腺癌分期指南摘要作为REK,比较三组表现:含REK且启用RAG(NotebookLM)、含REK但禁用RAG(Gemini 2.0 Flash)、不含REK且禁用RAG(Gemini 2.0 Flash)。基于CT影像对100个虚构胰腺癌病例进行分期,评估标准包括TNM分类、局部侵犯因素及可切除性分类。在含REK且启用RAG组中,量化了检索到的REK片段的充分性。结果:含REK且启用RAG组的分期准确率为70%,优于含REK但禁用RAG组(38%)和不含REK且禁用RAG组(35%)。在TNM分类上,前者准确率达80%,高于后者(55%)和对照组(50%)。此外,该组能明确呈现检索到的REK片段,检索准确率达92%。结论:NotebookLM(RAG-LLM)在胰腺癌分期任务中优于其内部模型Gemini 2.0 Flash,表明RAG有助于提升模型分期准确性。同时,其可展示引用来源的能力增强了医生可解释性,凸显其在临床诊断与分类中的应用潜力。
原文摘要 · Abstract (English)
Purpose: Retrieval-augmented generation (RAG) is a technology to enhance the functionality and reliability of large language models (LLMs) by retrieving relevant information from reliable external knowledge (REK). RAG has gained interest in radiology, and we previously reported the utility of NotebookLM, an LLM with RAG (RAG-LLM), for lung cancer staging. However, since the comparator LLM differed from NotebookLM's internal model, it remained unclear whether its advantage stemmed from RAG or inherent model differences. To better isolate RAG's impact and assess its utility across different cancers, we compared NotebookLM with its internal LLM, Gemini 2.0 Flash, in a pancreatic cancer staging experiment. Materials and Methods: A summary of Japan's pancreatic cancer staging guidelines was used as REK. We compared three groups - REK+/RAG+ (NotebookLM with REK), REK+/RAG- (Gemini 2.0 Flash with REK), and REK-/RAG- (Gemini 2.0 Flash without REK) - in staging 100 fictional pancreatic cancer cases based on CT findings. Staging criteria included TNM classification, local invasion factors, and resectability classification. In REK+/RAG+, retrieval accuracy was quantified based on the sufficiency of retrieved REK excerpts. Results: REK+/RAG+ achieved a staging accuracy of 70%, outperforming REK+/RAG- (38%) and REK-/RAG- (35%). For TNM classification, REK+/RAG+ attained 80% accuracy, exceeding REK+/RAG- (55%) and REK-/RAG- (50%). Additionally, REK+/RAG+ explicitly presented retrieved REK excerpts, achieving a retrieval accuracy of 92%. Conclusion: NotebookLM, a RAG-LLM, outperformed its internal LLM, Gemini 2.0 Flash, in a pancreatic cancer staging experiment, suggesting that RAG may improve LLM's staging accuracy. Furthermore, its ability to retrieve and present REK excerpts provides transparency for physicians, highlighting its applicability for clinical diagnosis and classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。