arXiv:2512.14640cs.CVcs.AI2025-12被引 1

首个多中心淋巴瘤分型基准,验证了深度学习在常规病理切片上精准分型的潜力与局限。

A Multicenter Benchmark of Multiple Instance Learning Models for Lymphoma Subtyping from HE-stained Whole Slide Images

  • 基于多实例学习框架,融合5个公开病理大模型与注意力/变压器聚合器。
  • 在同分布数据上准确率超80%,40倍分辨率已足够,更高分辨率无提升。
  • 跨中心测试性能降至约60%,凸显模型泛化难题,适合病理AI研究者参考。

及时准确的淋巴瘤诊断对指导治疗至关重要。标准诊断依赖于HE染色全片扫描图像结合免疫组化、流式细胞术和分子遗传学检测,需昂贵设备与专业人员,常导致治疗延迟。深度学习可直接从常规HE染色切片中提取诊断信息辅助病理科医生,但缺乏在多中心数据上的系统性基准。本文提出首个多中心淋巴瘤分型基准,涵盖四种常见亚型及健康对照组织。系统评估了五个公开病理基础模型(H-optimus-1, H0-mini, Virchow2, UNI2, Titan)与基于注意力(AB-MIL)和变压器(TransMIL)的多实例学习聚合器,在三个倍率(10x, 20x, 40x)下的表现。在同分布测试集上,模型多分类平衡准确率均超过80%,基础模型表现相当,聚合方法效果相近。倍率分析表明,40x分辨率已足够,更高分辨率或跨倍率聚合无性能提升。但在跨分布测试集上,性能显著下降至约60%,揭示明显的泛化挑战。为推动该领域发展,亟需覆盖更多罕见亚型的大规模多中心研究。我们提供了自动化基准测试流程,以支持未来工作。论文代码公开于 https://github.com/RaoUmer/LymphomaMIL。

原文摘要 · Abstract (English)

Timely and accurate lymphoma diagnosis is essential for guiding cancer treatment. Standard diagnostic practice combines hematoxylin and eosin (HE)-stained whole slide images with immunohistochemistry, flow cytometry, and molecular genetic tests to determine lymphoma subtypes, a process requiring costly equipment, and skilled personnel, causing treatment delays. Deep learning methods could assist pathologists by extracting diagnostic information from routinely available HE-stained slides directly, yet comprehensive benchmarks for lymphoma subtyping on multicenter data are lacking. In this work, we present the first multicenter lymphoma benchmark, covering four common lymphoma subtypes and healthy control tissue. We systematically evaluate five publicly available pathology foundation models (H-optimus-1, H0-mini, Virchow2, UNI2, Titan) combined with attention-based (AB-MIL) and transformer-based (TransMIL) multiple instance learning aggregators across three magnifications (10x, 20x, 40x). On in-distribution test sets, models achieve multiclass balanced accuracies exceeding 80% across all magnifications, with foundation models performing similarly, and aggregation methods showing comparable results. The magnification study reveals that 40x resolution is sufficient, with no performance gains from higher resolutions or cross-magnification aggregation. However, on out-of-distribution test sets, performance drops substantially to around 60%, highlighting significant generalization challenges. To advance the field, larger multicenter studies covering additional rare lymphoma subtypes are needed. We provide an automated benchmarking pipeline to facilitate such future research. Our paper codes is publicly available at https://github.com/RaoUmer/LymphomaMIL.

病理图像多实例学习淋巴瘤分型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。