整合37,000种化合物转录组数据,提升药物效应预测准确性
Chem-PerturBridge: a harmonized compendium of small molecule perturbation transcriptomic effects
- 统一多源数据格式与元信息,构建标准化小分子扰动资源
- 跨数据集相同化合物响应一致性弱,方向一致率高于幅度一致
- 作为预训练数据显著提升化合物表征性能,适合药物发现研究
大型扰动模型需要涵盖化学、细胞和检测技术多样性的训练数据。然而,现有小分子转录组资源在技术、元数据规范、对照组、剂量及预处理流程上分散割裂。我们提出 Chem-PerturBridge,一个整合多数据集的资源,包含超过 37,000 种化合物、136 个细胞背景和 125 万份转录组样本,覆盖八种检测类型,具有标准化标识符、元数据和可识别重复的条件级效应。我们利用该资源评估了跨数据集匹配条件的一致性以及数据集内重复的一致性。结果显示,多数数据集对相同化合物的精细对数倍数变化(logFC)排名与幅度一致性较弱,常低于同背景不同化合物的基线水平;而 logFC 方向一致性则更稳定,通常优于基线。进一步将 Chem-PerturBridge 用于化合物表征学习的预训练,在化合物留出的 OP3 评估划分下,其嵌入表示在各项指标上均优于仅基于 L1000、Morgan 指纹或无描述符的 OP3 基线。在 11 个数据集上的广泛分子留出评估中,基于 Chem-PerturBridge 训练的模型表现优于或相当于未使用该资源的模型。因此,Chem-PerturBridge 既可用于诊断跨数据集信号一致性,也支持异构扰动转录组数据的模型复用。
原文摘要 · Abstract (English)
Large perturbation models require training data encompassing chemical, cellular, and assay diversity. Current transcriptomic resources for small-molecule modeling, however, are fragmented across technologies, metadata conventions, controls, doses, and preprocessing pipelines. We introduce Chem-PerturBridge, a harmonized multi-dataset resource comprising over 37k compounds, 136 cellular contexts, and 1.25M transcriptomic samples across eight assay types, with standardized identifiers, metadata, and replicate-aware condition-level effects. We use the resource to evaluate matched-condition agreement across datasets and replicate agreement within datasets. Matched same-compound conditions generally show weak agreement in fine-grained logFC rankings and magnitudes across most dataset pairs, often falling below same-context different-compound baselines. In contrast, logFC direction agreement is substantially more stable and usually exceeds these baselines. We further evaluate Chem-PerturBridge as a pretraining resource for compound representation learning. Under a compound-held-out OP3 evaluation split, embeddings pretrained on Chem-PerturBridge improve over L1000-only embeddings, Morgan fingerprints, and the descriptor-free OP3 baseline across metrics. An extensive molecule-holdout evaluation across 11 datasets further shows that models trained on Chem-PerturBridge outperform or match those that are not. Chem-PerturBridge therefore supports both diagnostic evaluation of cross-dataset signature agreement and model-oriented reuse of heterogeneous perturbation transcriptomic data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。