arXiv:2609.02118cs.CV2026-09

提出新框架,让病理图像、基因和病历协同学习,提升诊断能力。

Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology

论文配图:Synergistic Information Disentanglement for Omni-modal Slide Representation Learning in Computational Pathology
图 1 · 摘自论文原文
  • 基于部分信息分解理论,分离多模态间高阶协同信号。
  • 在乳腺癌和肺癌数据集上预训练后,少样本任务表现超越现有方法。
  • 适合需要跨模态融合的医学影像分析研究者使用。

在计算病理学中,构建融合组织切片、基因组和临床报告的全模态自监督学习模型,有助于实现全幻灯片图像(WSIs)的可迁移表征学习。现有方法通过对比对齐将异构模态强制嵌入统一潜在空间,导致模态坍缩,使独特的协同诊断信号(记为$Φ$)被忽略,转而保留冗余信息。我们假设最强的无任务自监督训练信号来自提炼跨模态间的协同交互,而非仅对齐共享冗余。为此,提出 extsc{$Φ$-Omni}框架,基于部分信息分解(PID)理论,采用由$Φ$ID目标调控的协同信息瓶颈(SIB),显式抑制边际冗余并最大化不可约协同,从而提取高阶跨模态交互。在乳腺癌(n=1031)和肺癌(n=919)队列上预训练后, extsc{$Φ$-Omni}在五个独立外部数据集、八个任务上均展现出优于监督及自监督基线的少样本性能。源代码已公开。

原文摘要 · Abstract (English)

In computational pathology (CPath), developing omni-modal self-supervised learning (SSL) models that integrate histology, genomics, and clinical reports enables transferable representation learning for whole slide images (WSIs). Existing approaches implicitly force heterogeneous modalities into a uniform latent space by contrastive alignment, causing modality collapse where unique, synergistic diagnostic signals (termed as $\mathrm{\Phi}$) are discarded in favor of trivial redundancy. We hypothesize that the strongest task-agnostic SSL training signal stems from distilling the synergistic interactions over merely aligning shared redundancy. To this end, we introduce \textsc{$\mathrm{\Phi}$-Omni}, a synergistic information disentanglement framework grounded in Partial Information Decomposition (PID) theory for slide representation learning. Unlike standard contrastive approaches, \textsc{$\mathrm{\Phi}$-Omni} employs a Synergistic Information Bottleneck (SIB) regulated by the proposed $\mathrm{\Phi}\text{ID}$ objective, which explicitly suppresses marginal redundancy while maximizing irreducible synergy, thereby distilling high-order cross-modal interactions. Following pretraining on breast ($n$=1031) and lung ($n$=919) cohorts, \textsc{$\mathrm{\Phi}$-Omni} demonstrates superior few-shot performance across five independent external datasets spanning eight tasks compared to supervised and SSL baselines. Source code is available here.

病理图像多模态学习自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。