arXiv:2608.30420cs.CVcs.AI2026-09

提出新方法提升病理切片图像少样本标注精度,更贴近真实诊断场景。

Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols

论文配图:Whole-Slide Image Analysis under Realistic Few-Shot Annotation Protocols
图 1 · 摘自论文原文
  • 用条件随机场融合空间与生物学信息,处理切片中缺失类别问题。
  • 在四个数据集上,1次点击使宏平均F1提升24.2%,16次点击提升37.5%。
  • 设计模拟病理科医生操作的局部点击/涂鸦标注协议,更符合实际工作流。

全切片图像分析对癌症诊断具有重要临床价值,但依赖视觉-语言模型进行像素级零样本预测时噪声较大,需少量标注修正。当前的少样本传播方法多基于独立采样的切片块、平衡数据集和随机标注,忽略组织复杂结构、类别严重不平衡及真实标注行为。为此,本文提出SlideCRF,通过结合空间与生物学线索,并支持某些类别在单张切片中缺失的情况,改进条件随机场在全切片图像中的应用。同时,构建基于空间定位点击与涂鸦的真实标注协议,模拟病理医生逐步纠错等交互方式。在四个数据集上验证,使用1个或16个点击/类时,相较于零样本预测,宏平均F1分别提升24.2%和37.5%。

原文摘要 · Abstract (English)

Automating the analysis of whole-slide images has high clinical value, since characterizing cancers requires examining them in detail. Such analysis increasingly relies on vision-language models that provide patch-level zero-shot predictions. However, these predictions remain noisy and must be refined with a few annotations. A promising paradigm for this refinement is few-shot transduction. Rather than treating each patch independently, these methods leverage the relations between patches, together with a few annotations, to refine all predictions jointly. However, current transductive methods are evaluated under conditions that overlook key properties of whole-slide images: (i) datasets consist of independent patches extracted from multiple slides, ignoring the complex tissue organization; (ii) datasets are mostly balanced, whereas a single whole-slide image exhibits severe class imbalance, with several classes absent; and (iii) annotations are sampled at random, without reflecting how a pathologist annotates a limited number of regions. To align the transduction paradigm to realistic whole-slide settings, we introduce the following contributions. First, we propose SlideCRF, which adapts conditional random fields for whole-slide images by combining spatial and biological cues while accounting for classes that may be absent from a given slide. Second, we provide a set of realistic annotation protocols, based on spatially localized clicks and scribbles, modeling different pathologist interactions, such as the iterative correction of model errors. Across four datasets, we show that SlideCRF outperforms current transductive methods in macro F1, improving over the zero-shot predictions by +24.2% and +37.5% with one and 16 clicks per present class, respectively.

医学图像少样本学习病理分析条件随机场

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。