arXiv:2512.21331cs.CV2025-12被引 1

TICON通过上下文建模提升病理切片的局部表示,显著改善多种任务表现。

TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning

  • 基于Transformer的统一编码器,为任意病理切片模型生成上下文嵌入。
  • 在多个基准上达到新最好结果,包括HEST-Bench、THUNDER等。
  • 仅用1.1万张切片预训练,超越需35万张数据的现有模型。

小尺度病理切片图像的解读常需更大范围的上下文信息。我们提出TICON,一种基于Transformer的切片级别上下文建模器,可为计算病理学中任意应用生成丰富且具有上下文意义的嵌入。传统基于切片编码器的方法在剥离上下文的情况下提取嵌入,难以捕捉对局部与全局任务均至关重要的滑动切片级信息。此外,不同切片编码器在不同下游任务中表现各异。因此,亟需一个统一模型来整合并上下文化来自任意切片级基础模型的表示。TICON通过单一共享编码器,采用掩码建模目标进行预训练,实现多源切片编码器表示的统一与上下文化。实验表明,经TICON上下文化的嵌入在多个任务中显著提升性能,在切片级基准(如HEST-Bench、THUNDER、CATCH)和切片级基准(如Patho-Bench)上均取得新最佳结果。最后,我们在TICON基础上预训练一个聚合器,构建出滑动切片级基础模型,仅使用11,000张全切片图像,其性能超过需350,000张全切片图像预训练的现有最先进模型。

原文摘要 · Abstract (English)

The interpretation of small tiles in large whole slide images (WSI) often needs a larger image context. We introduce TICON, a transformer-based tile representation contextualizer that produces rich, contextualized embeddings for ''any'' application in computational pathology. Standard tile encoder-based pipelines, which extract embeddings of tiles stripped from their context, fail to model the rich slide-level information essential for both local and global tasks. Furthermore, different tile-encoders excel at different downstream tasks. Therefore, a unified model is needed to contextualize embeddings derived from ''any'' tile-level foundation model. TICON addresses this need with a single, shared encoder, pretrained using a masked modeling objective to simultaneously unify and contextualize representations from diverse tile-level pathology foundation models. Our experiments demonstrate that TICON-contextualized embeddings significantly improve performance across many different tasks, establishing new state-of-the-art results on tile-level benchmarks (i.e., HEST-Bench, THUNDER, CATCH) and slide-level benchmarks (i.e., Patho-Bench). Finally, we pretrain an aggregator on TICON to form a slide-level foundation model, using only 11K WSIs, outperforming SoTA slide-level foundation models pretrained with up to 350K WSIs.

病理图像上下文建模Transformer基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。