arXiv:2608.00105cs.CVq-bio.QM2026-08

用患者级验证发现病理大模型信号主要来自细胞计数而非几何结构。

What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer

论文配图:What Carries the Signal in Pathology Foundation-Model Atlases? A Patient-Level Controlled Benchmark in Breast Cancer
图 1 · 摘自论文原文
  • 以患者为单位评估模型,发现细胞计数特征贡献最大。
  • 嵌入向量预测分子程序得分达Spearman rho 0.25-0.56,显著优于随机。
  • 几何结构实际未带来增益,因拓扑由欧氏距离决定。

病理基础模型据称能从组织形态中编码分子程序,但证据多为群体层面的基因排序列表,而非对独立患者的预测。本研究以患者为分析单位,重建该分析:在285例有配对切片与RNA-seq的乳腺癌患者(TCGA-BRCA)上,使用11个冻结骨干网络、4个预设基因程序和分组交叉验证(按患者分组),均在折叠内完成所有预处理。对均值池化嵌入进行岭回归,对保留患者程序得分的预测达到Spearman相关系数0.25–0.56,其中UNI2在所有四个程序中表现最优(免疫程序0.556)。匹配置换检验显示每细胞原始p≈1e-4(10,000次置换),校正后p=0.0044。信号真实存在但并非全由形态决定。在相同患者与分组下,嵌入向量优于组织成分指标,对ER/腔面、增殖和免疫程序分别提升+0.280、+0.284、+0.479(p≤0.003),但在基底型中,仅分区分数达0.469,嵌入得分为0.493(p=0.77)。54个可解释的细胞计数特征在所有程序上接近嵌入性能(差距0.043–0.085)。几何机制无显著贡献,因测地线图通过欧氏最近邻搜索选邻点,仅重加权已选定边,拓扑本质仍为欧氏(黎曼减欧氏=+0.0010,95%CI [-0.0007, +0.0029])。一致应用几何结构反而更差(-0.0117)。岭回归优于图与度量解码器+0.097(95%CI [+0.069, +0.127])。文献中常见的驱动基因计数指标在此任务中几乎无信息量:91.8%的随机六基因面板能恢复≥5/6个驱动基因。

原文摘要 · Abstract (English)

Pathology foundation models are reported to encode molecular programmes in tissue morphology, but the evidence is usually a cohort-wide ranked gene list rather than a prediction for a held-out patient. We rebuild such an analysis with the patient as the unit of evidence and ask which pipeline component carries signal. Across 11 frozen backbones, four pre-specified gene programmes and 285 TCGA-BRCA patients with paired slides and RNA-seq (44 cells; GroupKFold by patient, all preprocessing fitted inside the fold), ridge regression on mean-pooled embeddings predicts held-out programme scores at Spearman rho = 0.25-0.56, UNI2 strongest on all four (immune 0.556). A matched permutation null gives raw p ~ 1e-4 at 10,000 permutations for every cell; Holm-adjusted p = 0.0044. The signal is real but not uniformly morphological. Against competing models on the same patients and folds, embeddings beat tissue composition for ER/luminal, proliferation and immune (+0.280, +0.284, +0.479; p <= 0.003) but not basal, where compartment fractions alone reach 0.469 against the embedding's 0.493 (p = 0.77). Fifty-four interpretable cell-count features come within 0.043-0.085 on every programme. The geometric machinery contributes nothing measurable, and we identify why: the geodesic graph selects neighbours by Euclidean nearest-neighbour search and only reweights edges already chosen, so the topology is Euclidean by construction (Riemannian minus Euclidean = +0.0010, 95% CI [-0.0007, +0.0029]). Applied consistently the geometry is worse (-0.0117). Ridge regression beats the graph-and-metric decoder by +0.097 (CI [+0.069, +0.127]). The driver-count metric common in this literature is near-uninformative here: 91.8% of random six-gene panels recover >=5/6 drivers.

病理模型细胞计数可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。