用精选综述+指南构建医学问答系统,精准回答长期新冠临床问题。
Demo: Guide-RAG: Evidence-Driven Corpus Curation for Retrieval-Augmented Generation in Long COVID
- 结合临床指南与高质量综述构建检索语料库
- 在专家题集上表现优于单一指南或海量文献库
- 适合医疗AI研发者与临床决策支持系统设计者
随着AI聊天机器人在临床医学中的应用增多,针对复杂新兴疾病构建有效框架面临挑战。我们开发并评估了六种用于长期新冠(LC)临床问答的检索增强生成(RAG)语料库配置,涵盖专家筛选来源至大规模文献数据库。评估采用大模型作为评判者,基于可信度、相关性与全面性指标,在新构建的专家生成临床问题数据集LongCOVID-CQ上进行测试。结果表明,融合临床指南与高质量系统综述的RAG配置,持续优于仅使用单一指南或大规模文献库的方法。研究发现,对于新兴疾病,基于精选二级文献的检索能实现共识文档与原始文献间的最佳平衡,既支持临床决策,又避免信息过载与过度简化。我们提出Guide-RAG,一个集成专家知识与全面文献库的聊天机器人系统及配套评估框架,可有效应答长期新冠临床问题。
原文摘要 · Abstract (English)
As AI chatbots gain adoption in clinical medicine, developing effective frameworks for complex, emerging diseases presents significant challenges. We developed and evaluated six Retrieval-Augmented Generation (RAG) corpus configurations for Long COVID (LC) clinical question answering, ranging from expert-curated sources to large-scale literature databases. Our evaluation employed an LLM-as-a-judge framework across faithfulness, relevance, and comprehensiveness metrics using LongCOVID-CQ, a novel dataset of expert-generated clinical questions. Our RAG corpus configuration combining clinical guidelines with high-quality systematic reviews consistently outperformed both narrow single-guideline approaches and large-scale literature databases. Our findings suggest that for emerging diseases, retrieval grounded in curated secondary reviews provides an optimal balance between narrow consensus documents and unfiltered primary literature, supporting clinical decision-making while avoiding information overload and oversimplified guidance. We propose Guide-RAG, a chatbot system and accompanying evaluation framework that integrates both curated expert knowledge and comprehensive literature databases to effectively answer LC clinical questions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。