arXiv:2606.29949eess.IVcs.AI2026-06

用少量数据让病理切片预测分子特征,无需重训练或测序。

Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction

论文配图:Data-Efficient Multimodal Alignment for Histopathology-based Molecular Prediction
图 1 · 摘自论文原文
  • 冻结基础模型,轻量对齐模块实现基因集查询预测通路活性
  • 多癌种数据集上检索效果提升25倍,形态相关通路预测准确率超0.5
  • 可直接从切片预测免疫微环境,适合临床研究与跨队列泛化

H&E染色全切片图像具有大规模可用性和丰富的空间上下文,但缺乏分子特异性;而批量RNA测序虽提供转录组范围分辨率,成本高且存档受限。我们证明,在冻结的病理学和RNA-seq基础模型之上训练轻量级对齐模块,可实现开放词汇的分子提示——通过基因集签名查询H&E切片以预测通路活性,无需测序或端到端重训练。在包含1,720例样本的多癌种队列上,对比基线方法检索性能提升25倍。系统分析揭示了渐进可预测性谱:基于形态的程序(如细胞周期、免疫相关)最易预测(R² > 0.5),无形态足迹的通路仍难以预测。在POSEIDON临床试验中验证其临床价值:H&E预测的鳞状细胞癌评分复现NSCLC亚型特征,预测的IFN-γ与PD-L1肿瘤细胞表达组一致。此外,描述免疫激活和纤维化的基因集可仅凭组织学预测已知的肿瘤微环境表型。进一步验证了该方法在未见队列中的泛化能力,并展示了数据高效领域自适应,建立了一种原生基于切片的分子分析框架。

原文摘要 · Abstract (English)

H&E-stained whole-slide images offer cohort-scale availability and rich spatial context but lack molecular specificity, whereas bulk RNA-seq provides transcriptome-wide resolution at high cost with limited archival availability. We show that training a lightweight alignment module atop frozen histopathology and RNA-Seq foundation models enables open-vocabulary molecular prompting -- querying H&E slides with gene-set signatures to predict pathway activity without sequencing or end-to-end retraining. Using contrastive learning on a multi-cancer cohort (N=1,720), we achieve a 25-fold improvement in retrieval over baseline methods. Systematic analysis reveals a graduated predictability spectrum: morphologically grounded programs (cell-cycle programs, immune-related) are most reliably predicted (R^2>0.5), while predicting pathways with no morphological footprint remains challenging as expected. We validate clinical utility on the POSEIDON clinical trial: H&E-predicted squamous cell carcinoma scores recapitulate NSCLC subtype identity and predicted IFN-gamma mirror PD-L1 tumor-cell expression groups. Furthermore, genesets describing immune activation and fibrosis predict known tumor microenvironment archetypes from histology alone. We further validate generalization of our approach across unseen cohorts and demonstrate data-efficient domain adaptation, establishing a slide-native framework for molecular analysis on H&E images.

病理分析分子预测多模态对齐数据高效

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。