arXiv:2510.08603cs.CL2025-10被引 4

构建病理领域RAG框架与评测集,提升医学大模型准确性与可信度。

YpathRAG:A Retrieval-Augmented Generation Framework and Benchmark for Pathology

  • 双通道检索融合密集与词表引导稀疏检索,增强病理文本召回。
  • 在挑战性问答集上平均准确率提升9.0%,最高达15.6%。
  • 适用于医疗AI研发、病理诊断辅助系统开发者。

大语言模型在通用任务中表现优异,但在病理等高门槛领域仍存在幻觉问题。现有方法依赖领域微调,无法扩展知识边界或强制证据约束。为此,我们构建了涵盖28个亚领域、包含153万段落的病理向量数据库,并提出面向病理的YpathRAG框架,采用BGE-M3密集检索与词表引导稀疏检索相结合的双通道混合检索策略,辅以基于LLM的支撑证据判断模块,实现检索-判断-生成闭环。同时发布两个评估基准:YpathR和YpathQA-M。在YpathR上,YpathRAG的Recall@5达到98.64%,相比基线提升23个百分点;在包含300个最挑战性问题的YpathQA-M上,通用与医学大模型的准确率平均提升9.0%,最高达15.6%。结果表明该框架显著提升了检索质量与事实可靠性,为病理导向RAG提供可扩展的构建范式与可解释的评估标准。

原文摘要 · Abstract (English)

Large language models (LLMs) excel on general tasks yet still hallucinate in high-barrier domains such as pathology. Prior work often relies on domain fine-tuning, which neither expands the knowledge boundary nor enforces evidence-grounded constraints. We therefore build a pathology vector database covering 28 subfields and 1.53 million paragraphs, and present YpathRAG, a pathology-oriented RAG framework with dual-channel hybrid retrieval (BGE-M3 dense retrieval coupled with vocabulary-guided sparse retrieval) and an LLM-based supportive-evidence judgment module that closes the retrieval-judgment-generation loop. We also release two evaluation benchmarks, YpathR and YpathQA-M. On YpathR, YpathRAG attains Recall@5 of 98.64%, a gain of 23 percentage points over the baseline; on YpathQA-M, a set of the 300 most challenging questions, it increases the accuracies of both general and medical LLMs by 9.0% on average and up to 15.6%. These results demonstrate improved retrieval quality and factual reliability, providing a scalable construction paradigm and interpretable evaluation for pathology-oriented RAG.

病理AIRAG医学大模型信息检索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。