用检索增强生成技术让AI更准地判断肺癌分期,还能指明依据来源。
Application of NotebookLM, a Large Language Model with Retrieval-Augmented Generation, for Lung Cancer Staging
- 引入外部权威指南作为知识源,让AI基于真实依据推理
- 在100个虚构病例中达到86%准确率,远超GPT-4o的25%-39%
- 能精准定位参考文献位置,帮助医生核查AI回答可靠性
目的:大型语言模型(LLM)如ChatGPT在放射科日益受到关注,但其临床应用的可靠性受幻觉和缺乏引用等问题制约。为解决这些问题,本研究聚焦最新技术——检索增强生成(RAG),使LLM可调用可靠外部知识(REK)。具体评估了新发布的具备RAG功能的LLM NotebookLM在肺癌分期中的应用价值。方法:将日本现行肺癌分期指南作为REK提供给NotebookLM,随后要求其根据CT影像对100个虚构病例进行分期,并评估准确性。对比实验使用金标准LLM GPT-4 Omni(GPT-4o),分别在有无REK条件下完成相同任务。结果:NotebookLM在肺癌分期中达到86%准确率,显著优于GPT-4o(有REK时39%,无REK时25%)。此外,NotebookLM在查找参考位置方面准确率达95%。结论:NotebookLM通过利用REK成功完成肺癌分期,性能优于GPT-4o,且能高精度定位参考来源,便于放射科医生验证其输出可靠性。本研究凸显了具备RAG能力的LLM在影像诊断中的潜力。
原文摘要 · Abstract (English)
Purpose: In radiology, large language models (LLMs), including ChatGPT, have recently gained attention, and their utility is being rapidly evaluated. However, concerns have emerged regarding their reliability in clinical applications due to limitations such as hallucinations and insufficient referencing. To address these issues, we focus on the latest technology, retrieval-augmented generation (RAG), which enables LLMs to reference reliable external knowledge (REK). Specifically, this study examines the utility and reliability of a recently released RAG-equipped LLM (RAG-LLM), NotebookLM, for staging lung cancer. Materials and methods: We summarized the current lung cancer staging guideline in Japan and provided this as REK to NotebookLM. We then tasked NotebookLM with staging 100 fictional lung cancer cases based on CT findings and evaluated its accuracy. For comparison, we performed the same task using a gold-standard LLM, GPT-4 Omni (GPT-4o), both with and without the REK. Results: NotebookLM achieved 86% diagnostic accuracy in the lung cancer staging experiment, outperforming GPT-4o, which recorded 39% accuracy with the REK and 25% without it. Moreover, NotebookLM demonstrated 95% accuracy in searching reference locations within the REK. Conclusion: NotebookLM successfully performed lung cancer staging by utilizing the REK, demonstrating superior performance compared to GPT-4o. Additionally, it provided highly accurate reference locations within the REK, allowing radiologists to efficiently evaluate the reliability of NotebookLM's responses and detect possible hallucinations. Overall, this study highlights the potential of NotebookLM, a RAG-LLM, in image diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。