arXiv:2603.17765q-bio.QMcs.AI2026-03被引 1

用相似病例增强报告生成,减少幻觉,提升临床可信度

Grounded Multimodal Retrieval-Augmented Drafting of Radiology Impressions Using Case-Based Similarity Search

  • 通过图像与文本嵌入融合检索相似病历
  • 检索召回率达0.95以上,显著优于纯图像检索
  • 适合需要可解释性与临床可信的医学AI应用

自动化放射科报告生成受到深度学习和大语言模型兴起的推动。然而,完全生成式方法常出现幻觉且缺乏临床依据,限制了其在真实工作流中的可靠性。本研究提出一种多模态检索增强生成(RAG)系统,用于胸部X光片报告的可信草稿生成。系统结合对比图像-文本嵌入、基于病例的相似性检索以及引用约束的草稿生成,确保生成内容与历史报告事实一致。使用MIMIC-CXR数据集的精选子集构建多模态检索数据库,图像嵌入由CLIP编码器生成,文本嵌入来自结构化报告片段。采用FAISS索引实现融合相似性框架,支持高效近邻检索。检索到的病例用于构建有依据的提示词,安全机制强制引用覆盖并基于置信度拒绝生成。实验表明,多模态融合显著提升检索性能,对临床相关发现的Recall@5超过0.95。该系统生成可解释、具引用溯源的输出,相比传统生成方法更具可信度。本工作展示了检索增强型多模态系统在可靠临床决策支持与放射科工作流辅助中的潜力。

原文摘要 · Abstract (English)

Automated radiology report generation has gained increasing attention with the rise of deep learning and large language models. However, fully generative approaches often suffer from hallucinations and lack clinical grounding, limiting their reliability in real-world workflows. In this study, we propose a multimodal retrieval-augmented generation (RAG) system for grounded drafting of chest radiograph impressions. The system combines contrastive image-text embeddings, case-based similarity retrieval, and citation-constrained draft generation to ensure factual alignment with historical radiology reports. A curated subset of the MIMIC-CXR dataset was used to construct a multimodal retrieval database. Image embeddings were generated using CLIP encoders, while textual embeddings were derived from structured impression sections. A fusion similarity framework was implemented using FAISS indexing for scalable nearest-neighbor retrieval. Retrieved cases were used to construct grounded prompts for draft impression generation, with safety mechanisms enforcing citation coverage and confidence-based refusal. Experimental results demonstrate that multimodal fusion significantly improves retrieval performance compared to image-only retrieval, achieving Recall@5 above 0.95 on clinically relevant findings. The grounded drafting pipeline produces interpretable outputs with explicit citation traceability, enabling improved trustworthiness compared to conventional generative approaches. This work highlights the potential of retrieval-augmented multimodal systems for reliable clinical decision support and radiology workflow augmentation

医学影像多模态检索增强报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。