用AI工具Sherpa Rx验证药基因组问答效果,提升用药决策精准度。
Validating Pharmacogenomics Generative Artificial Intelligence Query Prompts Using Retrieval-Augmented Generation (RAG)
- 结合CPIC与PharmGKB数据,用RAG增强生成准确回答。
- 在260个查询中准确率达4.9(5分制),完整度达4.8,召回率0.99。
- 比ChatGPT-4omini更优,适合临床医生做个性化用药参考。
本研究评估了基于大语言模型与检索增强生成(RAG)的AI工具Sherpa Rx在药基因组学领域的表现。该工具整合了临床药基因组学实施联盟(CPIC)指南与药基因组学知识库(PharmGKB)数据,生成上下文相关回应。使用涵盖26项CPIC指南的260个查询数据集,评估药物-基因相互作用、剂量建议及治疗意义。第一阶段仅嵌入CPIC数据,第二阶段额外加入PharmGKB内容。通过5分制量表评分准确率、相关性、清晰度、完整性,并计算召回率。比较了第一阶段与第二阶段、第二阶段与ChatGPT-4omini的准确性差异。20题实际应用测试显示,Sherpa Rx准确率达90%,显著优于其他模型。第二阶段虽未显著优于第一阶段,但显著优于ChatGPT-4omini。结果表明,融合多源数据库与RAG可有效提升AI在药基因组学中的性能。
原文摘要 · Abstract (English)
This study evaluated Sherpa Rx, an artificial intelligence tool leveraging large language models and retrieval-augmented generation (RAG) for pharmacogenomics, to validate its performance on key response metrics. Sherpa Rx integrated Clinical Pharmacogenetics Implementation Consortium (CPIC) guidelines with Pharmacogenomics Knowledgebase (PharmGKB) data to generate contextually relevant responses. A dataset (N=260 queries) spanning 26 CPIC guidelines was used to evaluate drug-gene interactions, dosing recommendations, and therapeutic implications. In Phase 1, only CPIC data was embedded. Phase 2 additionally incorporated PharmGKB content. Responses were scored on accuracy, relevance, clarity, completeness (5-point Likert scale), and recall. Wilcoxon signed-rank tests compared accuracy between Phase 1 and Phase 2, and between Phase 2 and ChatGPT-4omini. A 20-question quiz assessed the tool's real-world applicability against other models. In Phase 1 (N=260), Sherpa Rx demonstrated high performance of accuracy 4.9, relevance 5.0, clarity 5.0, completeness 4.8, and recall 0.99. The subset analysis (N=20) showed improvements in accuracy (4.6 vs. 4.4, Phase 2 vs. Phase 1 subset) and completeness (5.0 vs. 4.8). ChatGPT-4omini performed comparably in relevance (5.0) and clarity (4.9) but lagged in accuracy (3.9) and completeness (4.2). Differences in accuracy between Phase 1 and Phase 2 was not statistically significant. However, Phase 2 significantly outperformed ChatGPT-4omini. On the 20-question quiz, Sherpa Rx achieved 90% accuracy, outperforming other models. Integrating additional resources like CPIC and PharmGKB with RAG enhances AI accuracy and performance. This study highlights the transformative potential of generative AI like Sherpa Rx in pharmacogenomics, improving decision-making with accurate, personalized responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。