无需负样本即可解析生成模型,揭示抗体设计关键位点
Attribution assignment for deep-generative sequence models enables interpretability analysis using positive-only data
- 基于积分梯度构建生成模型可解释性方法GAMA
- 在已知真实特征的合成数据上验证了高准确性
- 适用于抗体等生物序列设计,无需负样本数据
生成式机器学习模型通过高效探索富含理想特性的生物序列空间,在药物设计中具有强大潜力。与需要正负样本的监督学习不同,生成模型(如LSTM)仅需正样本(例如高亲和力抗体)即可训练,这在负样本稀缺或定义不清的生物学场景中尤为有利。然而,生成模型缺乏有效的归因方法,阻碍了可解释性分析。为此,我们提出生成归因度量分析(GAMA),一种基于积分梯度的自回归生成模型归因方法。通过使用具有已知真实特征的合成数据评估GAMA的统计行为,并验证其恢复生物相关特征的能力。进一步在实验抗体-抗原结合数据上应用GAMA,证明其能实现模型可解释性并验证生成序列设计策略,且无需负样本数据。
原文摘要 · Abstract (English)
Generative machine learning models offer a powerful framework for therapeutic design by efficiently exploring large spaces of biological sequences enriched for desirable properties. Unlike supervised learning methods, which require both positive and negative labeled data, generative models such as LSTMs can be trained solely on positively labeled sequences, for example, high-affinity antibodies. This is particularly advantageous in biological settings where negative data are scarce, unreliable, or biologically ill-defined. However, the lack of attribution methods for generative models has hindered the ability to extract interpretable biological insights from such models. To address this gap, we developed Generative Attribution Metric Analysis (GAMA), an attribution method for autoregressive generative models based on Integrated Gradients. We assessed GAMA using synthetic datasets with known ground truths to characterize its statistical behavior and validate its ability to recover biologically relevant features. We further demonstrated the utility of GAMA by applying it to experimental antibody-antigen binding data. GAMA enables model interpretability and the validation of generative sequence design strategies without the need for negative training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。