用相似病例文本生成病理图像描述,减少幻觉和错误诊断。
Retrieval-Guided Generation for Safer Histopathology Image Captioning

- 从相似病例中检索专家文本,摘要生成描述,避免直接生成
- 在ARCH数据集上语义相似度达0.60,显著优于MedGemma的0.47
- 适合需要可审计、高可信度病理报告的临床场景
生成式视觉-语言模型虽能生成流畅的医学图像描述,但仍易出现幻觉、过度具体化诊断及事实不一致等严重问题。本文研究检索引导生成(RGG)作为更安全的替代方案:通过总结与图像视觉相似病例的专家文本来生成描述,而非从头生成。在ARCH病理图像数据集上,RGG的语义对齐度提升至约0.60(对比MedGemma的约0.47),且置信区间无重叠,表明性能提升显著。病理科医生主导的定性评估显示,该方法更好地保留了形态学相关术语,减少了未经支持的诊断;但也暴露出概念混淆和继承过度具体标签等失败模式。总体而言,检索引导生成提供更透明、可靠的路径,相比完全生成方法更具可审计性。
原文摘要 · Abstract (English)
Generative vision-language models can produce fluent medical image captions but remain prone to hallucination, over-specific diagnostic claims, and factual inconsistency-serious issues in pathology. We investigate retrieval-guided generation (RGG) as a safer alternative, where captions are formed by summarizing expert text from visually similar cases rather than generated de novo. On the ARCH histopathology dataset, RGG improves semantic alignment with ground truth, achieving cosine similarity of $\approx$0.60 versus $\approx$0.47 from MedGemma, with non-overlapping confidence intervals indicating a robust gain. A pathologist-led qualitative review shows better preservation of morphology-relevant terminology and fewer unsupported diagnoses, while revealing failure modes such as concept mixing and inherited over-specific labeling. Overall, retrieval-guided captioning offers a more transparent and reliable approach with clearer opportunities for auditing than fully generative methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。