用文本嵌入做先验,无监督推断基因调控网络
InfoSEM: A Deep Generative Model with Informative Priors for Gene Regulatory Network Inference
- 用基因文本嵌入作信息性先验,无需标签训练
- 在四个数据集上比现有模型提升38.5%
- 适合生物标志物发现场景,可揭示旧方法偏差
从基因表达数据推断基因调控网络(GRNs)对理解生物过程至关重要。现有监督模型虽性能高,但依赖昂贵的真值标签,易学习到基因特异性偏差(如真值互作的类别不平衡),而非真实调控机制。为此,我们提出InfoSEM,一种无监督生成模型,利用基因文本嵌入作为信息性先验,实现无需真值标签的GRN推断。当有标签数据时,可将其作为额外先验引入,避免偏差并进一步提升性能。此外,我们设计了一种生物合理性的评估框架,更贴近真实应用场景(如生物标志物发现),揭示了现有监督方法的隐含偏差。在四个数据集上,仅使用文本嵌入先验时,InfoSEM性能优于现有模型38.5%;加入标签先验后,性能再提升11.1%。
原文摘要 · Abstract (English)
Inferring Gene Regulatory Networks (GRNs) from gene expression data is crucial for understanding biological processes. While supervised models are reported to achieve high performance for this task, they rely on costly ground truth (GT) labels and risk learning gene-specific biases, such as class imbalances of GT interactions, rather than true regulatory mechanisms. To address these issues, we introduce InfoSEM, an unsupervised generative model that leverages textual gene embeddings as informative priors, improving GRN inference without GT labels. InfoSEM can also integrate GT labels as an additional prior when available, avoiding biases and further enhancing performance. Additionally, we propose a biologically motivated benchmarking framework that better reflects real-world applications such as biomarker discovery and reveals learned biases of existing supervised methods. InfoSEM outperforms existing models by 38.5% across four datasets using textual embeddings prior and further boosts performance by 11.1% when integrating labeled data as priors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。