arXiv:2608.16959eess.IVcs.CV2026-08中稿 · publication in the…

多倍率乳腺病理图像分类,提升准确率与可解释性。

MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology

论文配图:MagViT: Interpretable Multi-Magnification Transformers with Patient-Level Model Selection for Breast Histopathology
图 1 · 摘自论文原文
  • 多尺度融合+患者级模型选择,自适应整合不同放大倍数特征。
  • 在BreakHis上达96.43%患者准确率,跨数据集测试也表现稳健。
  • 可视化显示模型关注诊断关键区域,适合医学影像可信分析场景。

乳腺癌是全球女性中最常见的癌症之一。快速检测与早期治疗可阻止其进展至更复杂阶段并抑制转移。组织病理图像分类因其对细胞数据的强分析能力,成为癌症检测中最常见的任务。乳腺病理分类需同时处理多尺度组织形态及跨源域的临床相关泛化能力。本文提出MagViT,一种具有可解释性的多倍率变换器框架,包含尺度门控融合与患者级模型选择。该模型使用BreakHis四个放大倍数(40X、100X、200X、400X),以ViT为主干提取各尺度表征,并通过可学习门控机制融合,自动屏蔽缺失尺度。采用固定种子的患者级五折交叉验证,比较三种架构分支,最终选取患者级准确率最高且融合路径最简的分支作为最终模型。在BreakHis上,模型实现均值图像准确率0.9191、均值患者准确率0.9643、均值宏F1为0.9042。外部迁移实验在BUSI(图像准确率0.8306,宏F1 0.7480,患者准确率0.8291)和IDC(图像准确率0.8577,宏F1 0.8191,患者准确率0.8372)上表现出初步跨数据集泛化能力。Grad-CAM可视化显示模型聚焦于各倍率下的诊断意义区域。相较于以往基于ViT的BreakHis工作,本研究强调患者级选择与可复现协议下的跨数据集鲁棒性。

原文摘要 · Abstract (English)

Breast cancer is one of the most common types of cancer among women around the world. Rapid detection and early treatment can hinder its progress to more complex stages and can impede its spread to other parts of the body. Histopathological image classification is the most common task in cancer detection due to its robustness in analyzing cellular data. Breast histopathology classification requires handling both multi-scale tissue morphology and clinically relevant generalization beyond the source domain. This paper presents MagViT, an interpretable multi-magnification transformer framework with scale-gated fusion and patient-level model selection. The model uses four BreakHis magnifications (40X, 100X, 200X, 400X) and extracts per-scale representations with a ViT backbone, and combines them via a learnable gate that masks missing scales. Patient-level five-fold cross-validation with a fixed seed has been run and compared with three architectural branches. The most accurate branch is then selected as the final model due to the strongest patient-level accuracy while retaining the simplest fusion pathway. On BreakHis, our architecture achieves a mean image accuracy of 0.9191, a mean patient accuracy of 0.9643, and a mean macro-F1 of 0.9042. External transfer experiments provide preliminary evidence of cross-dataset generalization under controlled adaptation settings on BUSI (image accuracy 0.8306, macro-F1 0.7480, patient accuracy 0.8291) and IDC (image accuracy 0.8577, macro-F1 0.8191, patient accuracy 0.8372). Grad-CAM visualization indicates that the model focuses on diagnostically significant and meaningful regions across magnifications. Relative to prior ViT-centered BreakHis work, this study emphasizes patient-level selection and cross-dataset robustness under a reproducible protocol.

病理图像多尺度可解释性Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。