arXiv:2603.16185cs.LGcs.AI2026-03

通过分阶段学习,用少量患者数据高效适配药物反应模型。

Sample-Efficient Adaptation of Drug-Response Models to Patient Tumors under Strong Biological Domain Shift

  • 先无监督学习细胞和药物表征,再对齐标签进行适应。
  • 在仅有少量标注患者数据时,性能提升更快且更稳定。
  • 适合临床数据稀缺场景下的精准肿瘤治疗研究。

从临床前数据预测患者药物反应仍是精准肿瘤学的重大挑战,主要源于体外细胞系与患者肿瘤间的显著生物学差异。本文不追求提升体外预测绝对精度,而是探究将表征学习与任务监督显式分离是否能更高效地适应患者数据。提出一种分阶段迁移学习框架:首先利用大规模未标注药理基因组数据,通过基于自编码器的方法独立学习细胞与药物表征;随后在细胞系数据上对齐药物反应标签,并使用少量样本监督将模型适配至患者肿瘤。在涵盖域内、跨数据集及患者级设置的系统评估中发现,当源域与目标域重叠度高时,无监督预训练收益有限;但在适应患者肿瘤且标注数据极少的情况下,表现明显更优。该框架在少样本患者级适应中实现更快性能提升,同时在标准细胞系基准上保持与单阶段基线相当的准确性。结果表明,从无标注分子谱中学习结构化可迁移表征,可大幅减少临床监督所需数据量,为临床前到临床的高效转化提供可行路径。

原文摘要 · Abstract (English)

Predicting drug response in patients from preclinical data remains a major challenge in precision oncology due to the substantial biological gap between in vitro cell lines and patient tumors. Rather than aiming to improve absolute in vitro prediction accuracy, this work examines whether explicitly separating representation learning from task supervision enables more sample-efficient adaptation of drug-response models to patient data under strong biological domain shift. We propose a staged transfer-learning framework in which cellular and drug representations are first learned independently from large collections of unlabeled pharmacogenomic data using autoencoder-based representation learning. These representations are then aligned with drug-response labels on cell-line data and subsequently adapted to patient tumors using few-shot supervision. Through a systematic evaluation spanning in-domain, cross-dataset, and patient-level settings, we show that unsupervised pretraining provides limited benefit when source and target domains overlap substantially, but yields clear gains when adapting to patient tumors with very limited labeled data. In particular, the proposed framework achieves faster performance improvements during few-shot patient-level adaptation while maintaining comparable accuracy to single-phase baselines on standard cell-line benchmarks. Overall, these results demonstrate that learning structured and transferable representations from unlabeled molecular profiles can substantially reduce the amount of clinical supervision required for effective drug-response prediction, offering a practical pathway toward data-efficient preclinical-to-clinical translation.

药物反应少样本学习迁移学习精准肿瘤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。