用大模型实现病理图像中细胞分割与肿瘤微环境智能解读
SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology

- 基于双头结构的分割模型,结合大模型与梯度回归实现六类细胞精准分割
- 通过三阶段伪标签训练,在无标注数据上达到高质量分割效果
- 输出可读报告,适合临床研究与空间生物学分析
从常规H&E染色病理图像中解析肿瘤微环境(TME)需同步完成细胞分割、特征提取与可解释的临床报告生成。我们提出SegTME-UNI2统一框架,其核心为UNI2-UperHoVeR,一个将预训练于超百万张切片的ViT-Giant病理基础模型(UNI2-h)与两个并行的UperNet解码器结合的双头分割模型:一个用于六类语义分割,另一个通过水平-垂直梯度回归实现基于分水岭的核实例分离。针对大规模真实数据集缺乏像素级标注的问题,该模型采用三阶段渐进式伪标签课程训练,每阶段均从零开始训练新模型,仅靠伪标签质量提升推动性能迭代:第1阶段使用人工标注的PanNuke数据集(7,901张图像,189,744个核,0.25 μm/pixel);第2阶段使用第1阶段模型生成的熵过滤伪标签,覆盖271,711个TCGA-UT scale-0补丁(0.5 μm/pixel);第3阶段则利用第2阶段模型在全部1,608,060个跨六分辨率尺度(0.5–1.0 μm/pixel)的TCGA-UT补丁上生成伪标签。分割结果输入结构化特征提取管道,计算20余项每补丁的组成、形态、空间熵及细胞间距离指标,并编码为JSON传递至微调后的NVIDIA BioNeMo GPT模型,生成临床可读的TME描述。初步在预留的PanNuke与TCGA-UT分区上验证了框架可行性与内部一致性。伪标注的TCGA-UT数据集与UNI2-UperHoVeR检查点已公开,以支持大规模TME表征与空间生物学研究。
原文摘要 · Abstract (English)
Characterising the tumour microenvironment (TME) from routine H&E-stained histology images requires simultaneous cell segmentation, feature extraction, and interpretable clinical reporting. We present SegTME-UNI2, a unified framework addressing these requirements. Its core is UNI2-UperHoVeR, a dual-head segmentation model pairing the UNI2-h pathology foundation model (ViT-Giant, pretrained on >100M tiles from 100K slides) with two parallel UperNet decoders: one for six-class semantic segmentation and one for horizontal-vertical gradient regression enabling watershed-based nuclear instance separation. To address the lack of pixel-level annotations in large real-world repositories, UNI2-UperHoVeR undergoes a three-stage progressive pseudo-label curriculum. Each stage trains a fresh model without weight transfer, driving improvement entirely via increased pseudo-label quality: Stage 1: Uses human-annotated PanNuke (7,901 images, 189,744 nuclei, 0.25 um/pixel). Stage 2: Uses entropy-filtered pseudo-labels from the Stage 1 model on 271,711 TCGA-UT scale-0 patches (0.5 um/pixel). Stage 3: Uses pseudo-labels from the Stage 2 model on all 1,608,060 TCGA-UT patches across six resolution scales (0.5-1.0 um/pixel). Segmentation outputs feed a structured TME feature extraction pipeline computing 20+ per-patch compositional, morphological, spatial entropy, and intercellular distance metrics. These are encoded as JSON and passed to a fine-tuned NVIDIA BioNeMo GPT model to generate clinically interpretable TME narratives. Preliminary validation on held-out PanNuke and TCGA-UT partitions demonstrates framework feasibility and internal consistency. The pseudo-labelled TCGA-UT dataset and UNI2-UperHoVeR checkpoint are publicly released to support large-scale TME profiling and spatial biology research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。