arXiv:2505.04680cs.IR2025-05被引 7

用RAG提升医疗文档生成的准确与可信度,助力临床决策。

Retrieval Augmented Generation Evaluation for Health Documents

  • 构建RAGEv管道,融合先进实践确保医疗文本分析安全可靠。
  • 在短问答和长文本生成上均获高分,验证了RAG有效性。
  • 适合政策支持、科研辅助等需要精准知识提取的场景。

大型语言模型(LLM)在处理医疗文档和科学论文中的安全可信应用,可帮助临床医生、科学家和政策制定者克服信息过载,聚焦关键信息。检索增强生成(RAG)是一种有望在提升LLM潜力的同时增强输出准确性的方法。本报告评估了该方法在健康领域各类文档自动知识合成中的潜力与局限。为此,提出:(1) 自主开发的概念验证流程RAGEv(Retrieval Augmented Generation Evaluation),采用前沿实践实现医疗文档与科学论文的安全可信分析;(2) 针对基于LLM的文档检索与生成的评估工具集;(3) 用于验证结果准确性与真实性的基准数据集RAGEv-Bench。结果表明,谨慎实施RAG技术可有效缓解医疗文档处理中常见的LLM问题,在短答案(是/否)和长答案任务中均取得优异表现。其在日常政策支持工作中具有高度应用潜力,但仍需进一步努力以实现稳定可信的工具化。

原文摘要 · Abstract (English)

Safe and trustworthy use of Large Language Models (LLM) in the processing of healthcare documents and scientific papers could substantially help clinicians, scientists and policymakers in overcoming information overload and focusing on the most relevant information at a given moment. Retrieval Augmented Generation (RAG) is a promising method to leverage the potential of LLMs while enhancing the accuracy of their outcomes. This report assesses the potentials and shortcomings of such approaches in the automatic knowledge synthesis of different types of documents in the health domain. To this end, it describes: (1) an internally developed proof of concept pipeline that employs state-of-the-art practices to deliver safe and trustable analysis for healthcare documents and scientific papers called RAGEv (Retrieval Augmented Generation Evaluation); (2) a set of evaluation tools for LLM-based document retrieval and generation; (3) a benchmark dataset to verify the accuracy and veracity of the results called RAGEv-Bench. It concludes that careful implementations of RAG techniques could minimize most of the common problems in the use of LLMs for document processing in the health domain, obtaining very high scores both on short yes/no answers and long answers. There is a high potential for incorporating it into the day-to-day work of policy support tasks, but additional efforts are required to obtain a consistent and trustworthy tool.

医疗AIRAG知识提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。