从蛋白中无监督发现功能亚结构,提升功能预测精度。
BioBlobs: Unsupervised Discovery of Functional Substructures for Protein Function Prediction
- 将蛋白压缩为少量连贯亚结构(blobs),直接预测功能
- 仅用少量残基即达强基线性能,适配多种编码器
- 无需标注即可发现催化位点,适合大规模功能注释
蛋白质功能由局部的协同亚结构驱动,如催化三联体、结合口袋和结构模体,这些结构仅占蛋白残基的一小部分。现有基于蛋白编码器的流程未在亚结构层面建模,无法回答核心生物学问题:蛋白的哪个亚结构决定其功能?我们提出BioBlobs,一种编码器无关、端到端可微的框架,将蛋白压缩为一组连贯亚结构(blobs),并仅基于这些blobs预测功能,使每个blob对应一个候选功能区域。在多种蛋白功能预测任务及序列与结构基编码器上,BioBlobs表现匹配或优于强基线,且仅使用极少数残基。发现的blobs能自适应空间尺度,从局部催化位点到整个结构域。仅使用蛋白级标签训练,BioBlobs成功恢复了M-CSA数据库中的实验注释催化位点,验证了无监督功能亚结构发现能力,为未注释蛋白质组的大规模功能位点发现开辟路径。
原文摘要 · Abstract (English)
Protein function is driven by cohesive substructures, such as catalytic triads, binding pockets, and structural motifs, that occupy only a small fraction of a protein's residues. Yet existing pipelines built on protein encoders do not model proteins at the substructure level, leaving the central biological question unanswered: which substructure of a protein is responsible for its function? We introduce BioBlobs, an encoder-agnostic, end-to-end differentiable framework that compresses a protein into a small set of cohesive substructures (blobs) and predicts function from these blobs alone, so that each blob corresponds to a candidate functional region. Across diverse protein function prediction tasks and multiple sequence- and structure-based encoders, BioBlobs matches or exceeds strong baselines while operating on only a small fraction of residues. The discovered blobs adapt their spatial scale to the task, ranging from local catalytic sites to entire structural domains. Trained only on protein-level labels, BioBlobs recovers experimentally annotated catalytic sites in the M-CSA database, demonstrating unsupervised functional substructure discovery and opening a path to large-scale functional site discovery across the unannotated proteome.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。