arXiv:2508.18567cs.LGq-bio.QM2025-08被引 7

用稀疏自编码器从少量蛋白数据中精准预测功能并设计新蛋白

Sparse Autoencoders for Low-$N$ Protein Function Prediction and Design

  • 用稀疏自编码器分解蛋白质语言模型的嵌入,提取可解释的潜在特征
  • 仅需24个序列就超越基础模型,在低数据场景下表现更优
  • 通过调控潜在变量,83%情况下设计出高功能蛋白,适合生物设计应用

从氨基酸序列预测蛋白功能在数据稀缺(低-$N$)场景下仍是核心挑战,限制了机器学习驱动的蛋白设计。蛋白质语言模型(pLMs)通过提供进化信息嵌入,而稀疏自编码器(SAEs)能将这些嵌入分解为可解释的潜在变量,捕捉结构与功能特征。然而,SAEs在低-$N$功能预测与蛋白设计中的有效性尚未系统研究。本文评估了在微调后的ESM2嵌入上训练的SAEs,在多种适应性外推和蛋白工程任务中的表现。结果表明,仅用24个序列,SAEs在功能预测上持续优于或媲美原始ESM2基线,说明其稀疏潜在空间编码了紧凑且生物学意义明确的表示,能从有限数据中更好泛化。此外,通过调控预测性潜在变量,利用pLM表示中的生物基序,83%情况下生成了最高功能变体,显著优于仅使用ESM2的设计方法。

原文摘要 · Abstract (English)

Predicting protein function from amino acid sequence remains a central challenge in data-scarce (low-$N$) regimes, limiting machine learning-guided protein design when only small amounts of assay-labeled sequence-function data are available. Protein language models (pLMs) have advanced the field by providing evolutionary-informed embeddings and sparse autoencoders (SAEs) have enabled decomposition of these embeddings into interpretable latent variables that capture structural and functional features. However, the effectiveness of SAEs for low-$N$ function prediction and protein design has not been systematically studied. Herein, we evaluate SAEs trained on fine-tuned ESM2 embeddings across diverse fitness extrapolation and protein engineering tasks. We show that SAEs, with as few as 24 sequences, consistently outperform or compete with their ESM2 baselines in fitness prediction, indicating that their sparse latent space encodes compact and biologically meaningful representations that generalize more effectively from limited data. Moreover, steering predictive latents exploits biological motifs in pLM representations, yielding top-fitness variants in 83% of cases compared to designing with ESM2 alone.

蛋白设计稀疏编码小样本学习语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。