用大模型一次分析多张病理切片,生成精准报告。
PolyPath: Adapting a Large Multimodal Model for Multi-slide Pathology Report Generation
- 利用大模型长上下文能力,整合多张高倍率切片信息。
- 对最多5张切片的病例,生成报告临床准确率超68%。
- 适合需要跨切片综合诊断的病理医生和医学AI研究者。
组织病理学评估是医学诊断与治疗决策的关键,通常需整合多个切片的发现。现有计算病理学中的视觉语言能力多局限于小区域、低倍率区域或单张全切片图像(WSI),难以处理跨越多个高倍率区域的多张切片。本文利用具备100万词元上下文窗口的Gemini 1.5 Flash大模型,成功从最多40,000个768×768像素的图像块(对应10X倍率)中生成最终诊断报告,相当于长达11小时的视频内容。专家病理医生评估显示,在最多含5张切片的病例中,生成报告的临床准确性达到或优于原始报告的比例为68%(95%置信区间:[60%, 76%])。尽管在6张及以上切片时性能下降,但本研究证实了现代大模型长上下文能力在复杂病理报告生成任务中的巨大潜力。
原文摘要 · Abstract (English)
The interpretation of histopathology cases underlies many important diagnostic and treatment decisions in medicine. Notably, this process typically requires pathologists to integrate and summarize findings across multiple slides per case. Existing vision-language capabilities in computational pathology have so far been largely limited to small regions of interest, larger regions at low magnification, or single whole-slide images (WSIs). This limits interpretation of findings that span multiple high-magnification regions across multiple WSIs. By making use of Gemini 1.5 Flash, a large multimodal model (LMM) with a 1-million token context window, we demonstrate the ability to generate bottom-line diagnoses from up to 40,000 768x768 pixel image patches from multiple WSIs at 10X magnification. This is the equivalent of up to 11 hours of video at 1 fps. Expert pathologist evaluations demonstrate that the generated report text is clinically accurate and equivalent to or preferred over the original reporting for 68% (95% CI: [60%, 76%]) of multi-slide examples with up to 5 slides. While performance decreased for examples with 6 or more slides, this study demonstrates the promise of leveraging the long-context capabilities of modern LMMs for the uniquely challenging task of medical report generation where each case can contain thousands of image patches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。