arXiv:2604.22832cs.CVcs.AI2026-04

用药物基因表达指导图像学习,提升新药预测能力

Intervention-Aware Multiscale Representation Learning from Imaging Phenomics and Perturbation Transcriptomics

论文配图:Intervention-Aware Multiscale Representation Learning from Imaging Phenomics and Perturbation Transcriptomics
图 1 · 摘自论文原文
  • 用转录组信息生成软标签,指导显微图像表示学习
  • 在新药和靶点发现任务上,比基线方法提升显著
  • 适合药物研发中缺乏完整配对数据的场景

基于显微镜的表型分析适用于药物发现,但缺乏转录组学的机制深度,而转录组学成本高且数据稀缺。现有跨模态方法或仅用图像辅助其他模态,或简单按样本身份对齐表示,忽略细胞类型与剂量差异,在弱配对数据下限制泛化能力。本文提出一种干预感知蒸馏框架,利用扰动转录组学引导图像表示学习。一个转录组条件下的教师模型整合基因表达与干预元信息,生成基于药物相似性的化学感知码本中的软分布;教师使用微调的单细胞基础模型编码细胞类型上下文并解耦剂量效应。仅含图像的学生模型从显微图像中学习预测这些分布,蒸馏机制知识,测试时可独立运行。该设计强调干预语义而非身份对齐,显式处理剂量与细胞类型不匹配问题。我们提供理论保证,证明转录组指导可收紧图像预测的风险界。在与L1000配对的Cell Painting和RxRx数据集上,本方法在未见干预的一次性迁移和药物靶点基因发现上均显著优于自监督与对齐基线。

原文摘要 · Abstract (English)

Microscopy-based phenotypic profiling is scalable for drug discovery but lacks the mechanistic depth of transcriptomics, which remains costly and scarce. Existing multimodal approaches either use images to support other modalities or naively align representations by sample identity, ignoring cell-type and dose variations in weakly paired data-limiting generalization to unseen interventions. In this paper, we introduce an intervention-aware distillation framework that leverages perturbational transcriptomics to guide image representation learning. A transcriptome-conditioned teacher integrates gene expression and intervention metadata to produce soft distributions over a chemistry-aware codebook organized by drug similarity. The teacher employs a fine-tuned single-cell foundation model to encode cell-type context and disentangle dose effects. An image-only student learns to predict these distributions from microscopy alone, distilling mechanistic knowledge while operating independently at test time. This design emphasizes intervention semantics rather than identity alignment and explicitly handles dose and cell-type mismatches. We provide theoretical guarantees showing that transcriptomic guidance tightens the risk bound for image-based prediction. On Cell Painting and RxRx datasets paired with L1000, our method significantly improves one-shot transfer to unseen interventions and drug-target gene discovery compared to self-supervised and alignment baselines.

多模态学习药物发现表型分析蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。