对比蛋白与自然语言模型的注意力分布差异,提出高效预测新方法。
Protein Language Models Diverge from Natural Language: Comparative Analysis and Improved Inference
- 通过分析注意力头信息分布,揭示蛋白语言模型运作机制差异。
- 引入早退机制,提升非结构属性预测准确率0.4~7.01个百分点。
- 兼顾效率与精度,适合生物序列分析任务快速推理需求。
现代蛋白质语言模型(PLMs)借鉴自然语言处理中的Transformer架构,用于预测蛋白质功能与性质。然而,蛋白质语言与自然语言存在本质差异:仅20种氨基酸构成却具有丰富的功能空间。这促使我们研究Transformer架构在蛋白质领域的不同表现,并探索如何更有效地利用PLMs解决蛋白质相关任务。本文首次直接比较了蛋白质与自然语言领域中注意力头信息在各层的分布差异。此外,我们改进了原本用于提升自然语言模型效率的早退技术,使其在蛋白质非结构属性预测中同时实现更高准确率与显著效率提升——准确率提高0.4至7.01个百分点,效率提升超10%。该研究为跨域语言模型行为对比开辟新方向,推动生物序列语言建模发展。
原文摘要 · Abstract (English)
Modern Protein Language Models (PLMs) apply transformer-based model architectures from natural language processing to biological sequences, predicting a variety of protein functions and properties. However, protein language has key differences from natural language, such as a rich functional space despite a vocabulary of only 20 amino acids. These differences motivate research into how transformer-based architectures operate differently in the protein domain and how we can better leverage PLMs to solve protein-related tasks. In this work, we begin by directly comparing how the distribution of information stored across layers of attention heads differs between the protein and natural language domain. Furthermore, we adapt a simple early-exit technique-originally used in the natural language domain to improve efficiency at the cost of performance-to achieve both increased accuracy and substantial efficiency gains in protein non-structural property prediction by allowing the model to automatically select protein representations from the intermediate layers of the PLMs for the specific task and protein at hand. We achieve performance gains ranging from 0.4 to 7.01 percentage points while simultaneously improving efficiency by over 10 percent across models and non-structural prediction tasks. Our work opens up an area of research directly comparing how language models change behavior when moved into the protein domain and advances language modeling in biological domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。