arXiv:2604.12100cs.CV2026-04

让病理图像分类同时看全局和局部,提升模型对癌症位置的判断能力。

PC-MIL: Decoupling Feature Resolution from Supervision Scale in Whole-Slide Learning

  • 用不同尺度的局部监督信号解耦特征分辨率与标签粒度
  • 在2mm尺度上锚定监督,使模型学会识别微米级肿瘤区域
  • 多尺度训练提升跨场景泛化性,适合临床病理分析场景

计算病理学中的全切片图像(WSI)分类通常采用仅基于全局标签的滑动窗口多重实例学习(MIL),但这种做法本质上约束不足:仅优化整体标签会促使模型忽略解剖结构,无法学习有意义的空间定位。临床医生关注的是毫米级区域内的肿瘤负荷、局灶病灶和组织架构模式,而标准MIL仅要求判断“某处是否有癌变”。为此,我们提出渐进式上下文多重实例学习(PC-MIL),将监督尺度作为核心设计维度,不改变放大倍数或补丁大小,而是通过固定20x特征,以毫米为单位调整MIL袋的范围,并在2mm临床尺度上锚定监督,以保留肿瘤负担一致性并避免尺度与病灶密度混淆。PC-MIL逐步混合全片与区域级监督,在可控比例下实现训练上下文与测试上下文的分析。在来自五个公开数据集的1,476例前列腺WSI上进行二分类癌症检测实验表明,解剖上下文是独立于特征分辨率的泛化轴线:适度的区域监督可提升跨上下文性能,均衡的多上下文训练在不牺牲全局性能的前提下,稳定了全片与区域评估下的准确率。结果证明监督范围塑造了MIL归纳偏置,支持基于解剖结构的全切片图像泛化。

原文摘要 · Abstract (English)

Whole-slide image (WSI) classification in computational pathology is commonly formulated as slide-level Multiple Instance Learning (MIL) with a single global bag representation. However, slide-level MIL is fundamentally underconstrained: optimizing only global labels encourages models to aggregate features without learning anatomically meaningful localization. This creates a mismatch between the scale of supervision and the scale of clinical reasoning. Clinicians assess tumor burden, focal lesions, and architectural patterns within millimeter-scale regions, whereas standard MIL is trained only to predict whether "somewhere in the slide there is cancer." As a result, the model's inductive bias effectively erases anatomical structure. We propose Progressive-Context MIL (PC-MIL), a framework that treats the spatial extent of supervision as a first-class design dimension. Rather than altering magnification, patch size, or introducing pixel-level segmentation, we decouple feature resolution from supervision scale. Using fixed 20x features, we vary MIL bag extent in millimeter units and anchor supervision at a clinically motivated 2mm scale to preserve comparable tumor burden and avoid confounding scale with lesion density. PC-MIL progressively mixes slide- and region-level supervision in controlled proportions, enabling explicit train-context x test-context analysis. On 1,476 prostate WSIs from five public datasets for binary cancer detection, we show that anatomical context is an independent axis of generalization in MIL, orthogonal to feature resolution: modest regional supervision improves cross-context performance, and balanced multi-context training stabilizes accuracy across slide and regional evaluation without sacrificing global performance. These results demonstrate that supervision extent shapes MIL inductive bias and support anatomically grounded WSI generalization.

病理图像多实例学习解剖上下文监督尺度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。