用流匹配模型从病理切片预测全基因组表达,更准确且可解释。
RNA-FM: Flow-Matching Generative Model for Genome-wide RNA-Seq Prediction

- 将基因表达预测建模为连续时间条件传输,学习形态到表达的动态映射
- 在多个数据集上优于现有方法,尤其在跨组织和跨批次场景表现突出
- 支持通路级结构建模,适合生物医学研究与临床转化应用
病理全切片图像(WSI)在临床中常规获取,包含丰富的组织形态信息,但缺乏直接定义病理状态的分子架构与功能程序;而RNA测序(RNA-seq)虽能提供全基因组转录谱,成本较高。这促使了基于WSI的全基因组转录组预测需求。现有方法多依赖确定性回归的一一映射,难以捕捉生物异质性和预测不确定性。本文提出RNA-FM,一种用于从WSI预测全基因组批量RNA-seq的流匹配生成框架。RNA-FM将转录组预测建模为连续时间条件传输问题,学习一个速度场,将简单先验分布映射至在形态条件下的目标基因表达分布。通过整合通路级结构,RNA-FM实现了可扩展且生物学可解释的全基因组基因表达推断。大量实验表明,RNA-FM持续优于当前最优方法,同时保持生物学合理性。代码已公开于 https://github.com/YXSong000/RNA-FM。
原文摘要 · Abstract (English)
Histopathology whole-slide images (WSIs) are routinely acquired in clinical practice and contain rich tissue morphology but lack direct molecular architecture and functional programs defining pathological states, whereas RNA sequencing (RNA-seq) provides genome-wide transcriptional profiles at substantial cost, thereby motivating WSI-based genome-wide transcriptomic prediction. Existing approaches for predicting gene expression from WSIs predominantly rely on deterministic regression with one-to-one mapping, limiting their ability to capture biological heterogeneity and predictive uncertainty. We propose RNA-FM, a flow-matching generative framework for genome-wide bulk RNA-seq prediction from WSIs. RNA-FM formulates transcriptomic prediction as a continuous-time conditional transport problem, learning a velocity field that maps a simple prior to the target gene expression distribution conditioned on morphologies. By integrating pathway-level structure, RNA-FM enables scalable and biologically interpretable genome-wide gene expression imputation. Extensive experiments demonstrate that RNA-FM consistently outperforms state-of-the-art approaches while maintaining biological meaningfulness. Code is available at https://github.com/YXSong000/RNA-FM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。