arXiv:2412.01388cs.LG2024-12被引 3

用偏好优化提升蛋白语言模型,加速细胞治疗候选药物筛选。

Harnessing Preference Optimisation in Protein LMs for Hit Maturation in Cell Therapy

  • 用高通量实验数据微调蛋白语言模型,采用偏好学习策略。
  • 微调后模型预测结果与生物实验高度相关,支持少样本药物优化。
  • 适用于免疫治疗药物研发,可推广至其他疗法场景。

细胞和免疫治疗通过调节免疫系统,在治疗癌症和自身免疫疾病方面展现出变革性潜力。但这类疗法研发成本高昂,多数候选药物无法通过实验室测试。尽管机器学习在蛋白质工程中取得突破,其在免疫治疗中的应用仍受限于缺乏大规模标准化数据集以及细胞系统的复杂性。本文利用高通量实验平台生成可用于微调蛋白语言模型的数据,证明基于偏好任务微调的模型表现出与生物实验惊人的一致性,并可应用于嵌合抗原受体(CARs)的少样本候选药物优化。该概念验证为将机器学习引入免疫治疗开辟了新路径,且具备向其他治疗领域扩展的潜力。

原文摘要 · Abstract (English)

Cell and immunotherapy offer transformative potential for treating diseases like cancer and autoimmune disorders by modulating the immune system. The development of these therapies is resource-intensive, with the majority of drug candidates failing to progress beyond laboratory testing. While recent advances in machine learning have revolutionised areas such as protein engineering, applications in immunotherapy remain limited due to the scarcity of large-scale, standardised datasets and the complexity of cellular systems. In this work, we address these challenges by leveraging a high-throughput experimental platform to generate data suitable for fine-tuning protein language models. We demonstrate how models fine-tuned using a preference task show surprising correlations to biological assays, and how they can be leveraged for few-shot hit maturation in CARs. This proof-of-concept presents a novel pathway for applying ML to immunotherapy and could generalise to other therapeutic modalities.

蛋白质语言模型免疫治疗少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。