给蛋白语言模型的神经元自动打标签,让生成和解释更精准。
Automated Neuron Labelling Enables Generative Steering and Interpretability in Protein Language Models
- 用自然语言自动标注百万级神经元,基于生物特性描述。
- 通过激活引导生成目标蛋白,实现分子量、结构等属性精准控制。
- 揭示模型规模与神经元分布的规律,助力可解释性研究。
蛋白语言模型(PLMs)蕴含丰富生物学信息,但其内部神经元表征尚不清晰。本文提出首个自动化框架,为PLM中每个神经元赋予基于生物知识的自然语言描述。相比依赖稀疏自编码器或人工标注的旧方法,本方法可扩展至数十万神经元,揭示单个神经元对多种生化与结构特性的选择性敏感。我们进一步开发一种基于神经元激活的引导生成方法,能生成具有特定性质的蛋白质,实现分子量、不稳定指数等生化属性,以及α螺旋、锌指等二级与三级结构特征的收敛。最后,通过分析不同规模模型中标注神经元的分布,发现PLM的缩放规律及结构化的神经元空间分布。
原文摘要 · Abstract (English)
Protein language models (PLMs) encode rich biological information, yet their internal neuron representations are poorly understood. We introduce the first automated framework for labeling every neuron in a PLM with biologically grounded natural language descriptions. Unlike prior approaches relying on sparse autoencoders or manual annotation, our method scales to hundreds of thousands of neurons, revealing individual neurons are selectively sensitive to diverse biochemical and structural properties. We then develop a novel neuron activation-guided steering method to generate proteins with desired traits, enabling convergence to target biochemical properties like molecular weight and instability index as well as secondary and tertiary structural motifs, including alpha helices and canonical Zinc Fingers. We finally show that analysis of labeled neurons in different model sizes reveals PLM scaling laws and a structured neuron space distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。