arXiv:2607.17671cs.LG2026-07

根据细胞响应信号,精准匹配最可能的作用靶点和化合物。

GeneSpeak-FP: Target and Compound Retrieval from Observed Cell-Level Perturbation Signatures

  • 用Transformer模型将细胞扰动信号映射为靶点与分子向量,联合训练目标与结构-转录对齐。
  • 在379种化合物库中,靶点召回率@10达40.8%,化合物命中率@1为12.9%。
  • 首次实现单模型同时恢复已知靶点与化合物身份,适合药物重定位研究。

大规模单细胞扰动图谱使我们能提出逆问题:给定一个观测到的转录响应,哪些固定库中的注释靶点和化合物最符合该响应?我们提出 model,一种用于封闭库场景的Transformer检索模型。每个输入是一个由处理细胞与细胞系特异性DMSO对照均值对比形成的细胞级扰动信号。编码器将信号映射为靶点检索向量和分子嵌入向量,通过监督靶点损失和结构-转录对齐联合训练。我们在包含10,505个训练和1,168个验证药物-细胞系对的Tahoe-100M条件数据集上评估,采用化合物内分层90/10条件对划分。由于化合物和细胞系同时出现在两个分区中,实验衡量的是未见条件对的检索能力,而非对未见化合物或细胞背景的泛化能力。在38,400个采样验证细胞的蒙特卡洛评估中, model 达到靶点召回率@10为0.408,召回率@20为0.544,同时化合物命中率@1为0.129,命中率@10为0.343,平均倒数排名为0.205。独立诊断评估显示主模型结果与随机向量控制及事后基因袋模型相比有显著提升。结果表明,单一多任务模型可在评估的Tahoe-100M封闭库设置下,从观测到的细胞水平响应中恢复已标注靶点和记录的化合物身份。对未见化合物和细胞背景的泛化能力仍待验证。

原文摘要 · Abstract (English)

Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response? We present \model, a Transformer retrieval model for this closed-library setting. Each input is a cell-level perturbation signature formed by contrasting one treated cell with a cell-line-specific mean DMSO reference. The encoder maps the signature to a target-retrieval vector and a molecular-embedding vector, trained jointly with supervised target losses and structure--transcriptome alignment. We evaluate on Tahoe-100M conditions with mapped target annotations using a within-compound stratified 90/10 condition-pair split of 10,505 training and 1,168 validation drug--cell-line pairs. Because compounds and cell lines can occur in both partitions, the experiment measures held-out condition-pair retrieval rather than generalization to unseen compounds or cellular contexts. In a Monte Carlo evaluation over 38,400 sampled validation cells, \model\ achieved target Recall@10 of 0.408 and Recall@20 of 0.544, together with compound Hit@1 of 0.129, Hit@10 of 0.343, and mean reciprocal rank of 0.205 over a 379-compound bank. A separate diagnostic evaluation produced nearly identical values for the main model and large gains over a random-vector control and post-hoc bag-of-genes controls. These results demonstrate that a single multi-task model can recover both mapped target annotations and recorded compound identities from observed cell-level responses in the evaluated Tahoe-100M closed-library setting. Generalization to unseen compounds and cellular contexts remains to be established.

靶点发现药物重定位单细胞分析Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。