用可解释概念控制蛋白生成,性能不降反升
Concept Bottleneck Language Models For protein design
- 神经元对应可解释概念,通过干预概念值精准调控蛋白属性
- 概念值改变幅度是基线模型的3倍,生成效果更可控
- 适合需要透明决策过程的药物设计场景
我们提出概念瓶颈蛋白语言模型(CB-pLM),一种生成式掩码语言模型,其某层中每个神经元对应一个可解释的概念。该架构带来三大优势:(i) 控制性:可通过干预概念值精确调控生成蛋白属性,实现的期望概念值变化幅度是基线模型的3倍;(ii) 可解释性:概念值与预测词元间存在线性映射,可透明分析模型决策过程;(iii) 可调试性:透明结构便于模型调试。模型在预训练困惑度和下游任务表现上与传统掩码蛋白语言模型相当,证明可解释性不牺牲性能。尽管该方法适用于任何语言模型,我们聚焦于掩码蛋白语言模型,因其在药物发现中的重要性,且可通过真实实验和专家知识验证模型能力。我们训练了从2400万到30亿参数的CB-pLM,成为迄今最大规模的概念瓶颈模型,也是首个支持生成式语言建模的概念瓶颈模型。
原文摘要 · Abstract (English)
We introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our architecture offers three key benefits: i) Control: We can intervene on concept values to precisely control the properties of generated proteins, achieving a 3 times larger change in desired concept values compared to baselines. ii) Interpretability: A linear mapping between concept values and predicted tokens allows transparent analysis of the model's decision-making process. iii) Debugging: This transparency facilitates easy debugging of trained models. Our models achieve pre-training perplexity and downstream task performance comparable to traditional masked protein language models, demonstrating that interpretability does not compromise performance. While adaptable to any language model, we focus on masked protein language models due to their importance in drug discovery and the ability to validate our model's capabilities through real-world experiments and expert knowledge. We scale our CB-pLM from 24 million to 3 billion parameters, making them the largest Concept Bottleneck Models trained and the first capable of generative language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。