arXiv:2409.15574cs.CV2024-09被引 10

用多尺度病理图像生成临床级报告,提升医生效率。

Clinical-grade Multi-Organ Pathology Report Generation for Multi-scale Whole Slide Images via a Semantically Guided Medical Text Foundation Model

  • 基于多尺度区域视觉变压器提取图像特征,指导视觉语言模型训练
  • 在结肠和肾脏多器官数据集上达成0.68的METEOR分数
  • 无需人工标注即可自动生成报告,适合临床辅助诊断场景

视觉语言模型(VLM)在自然语言理解和图像识别任务中取得成功,但在全切片图像(WSI)病理报告生成中的应用受限于多尺度WSI的巨大尺寸和标注成本。现有研究大多缺乏临床有效性验证。为此,我们提出一种患者级多器官病理报告生成(PMPRG)模型,利用所提出的多尺度区域视觉变压器(MR-ViT)模型提取的多尺度WSI特征及其真实病理报告,引导VLM训练以实现精准报告生成。模型基于关键特征关注的区域特征自动生成报告。我们在包含结肠和肾脏等多器官的WSI数据集上评估该模型,获得0.68的METEOR得分,验证了方法的有效性。该模型可帮助病理医生高效生成涉及多个WSI的患者报告。

原文摘要 · Abstract (English)

Vision language models (VLM) have achieved success in both natural language comprehension and image recognition tasks. However, their use in pathology report generation for whole slide images (WSIs) is still limited due to the huge size of multi-scale WSIs and the high cost of WSI annotation. Moreover, in most of the existing research on pathology report generation, sufficient validation regarding clinical efficacy has not been conducted. Herein, we propose a novel Patient-level Multi-organ Pathology Report Generation (PMPRG) model, which utilizes the multi-scale WSI features from our proposed multi-scale regional vision transformer (MR-ViT) model and their real pathology reports to guide VLM training for accurate pathology report generation. The model then automatically generates a report based on the provided key features attended regional features. We assessed our model using a WSI dataset consisting of multiple organs, including the colon and kidney. Our model achieved a METEOR score of 0.68, demonstrating the effectiveness of our approach. This model allows pathologists to efficiently generate pathology reports for patients, regardless of the number of WSIs involved.

病理报告生成多尺度图像视觉语言模型临床辅助

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。