用检索增强生成技术,让体检报告解读更个性化。
Lab-AI: Using Retrieval Augmentation to Enhance Language Models for Personalized Lab Test Interpretation in Clinical Medicine
- 基于可信医学源的检索增强生成,动态获取患者专属参考范围。
- 因子检索F1达0.948,参考范围准确率高达0.995。
- 适合临床医生与患者门户,提升个性化医疗沟通效率。
准确解读检验结果对临床医学至关重要,但多数患者门户仍使用通用参考范围,忽略年龄、性别等条件因素。本研究提出Lab-AI,一个交互式系统,利用从可信健康资源中检索增强生成(RAG)技术,提供个性化参考范围。系统包含因子检索与参考范围检索两个模块。我们在122项检验上进行了测试:40项受条件影响,82项不受影响。对于有条件检验,参考范围依赖于患者具体信息。结果显示,GPT-4-turbo结合RAG在因子检索上达到0.948的F1分数,在参考范围检索上准确率达0.995。相比最优非RAG系统,因子检索性能提升33.5%;在问题级别和检验级别上,参考范围检索分别提升132%和100%。这些发现表明Lab-AI有潜力显著提升患者对检验结果的理解。
原文摘要 · Abstract (English)
Accurate interpretation of lab results is crucial in clinical medicine, yet most patient portals use universal normal ranges, ignoring conditional factors like age and gender. This study introduces Lab-AI, an interactive system that offers personalized normal ranges using retrieval-augmented generation (RAG) from credible health sources. Lab-AI has two modules: factor retrieval and normal range retrieval. We tested these on 122 lab tests: 40 with conditional factors and 82 without. For tests with factors, normal ranges depend on patient-specific information. Our results show GPT-4-turbo with RAG achieved a 0.948 F1 score for factor retrieval and 0.995 accuracy for normal range retrieval. GPT-4-turbo with RAG outperformed the best non-RAG system by 33.5% in factor retrieval and showed 132% and 100% improvements in question-level and lab-level performance, respectively, for normal range retrieval. These findings highlight Lab-AI's potential to enhance patient understanding of lab results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。