arXiv:2505.11812cs.LGcs.CL2025-05被引 6

首个针对蛋白质精细功能注释的大型基准,助力解析蛋白作用机制

VenusX: Unlocking Fine-Grained Functional Understanding of Proteins

  • 构建跨残基、片段、结构域的多层级功能注释任务
  • 涵盖超87万样本,支持同家族与跨家族评估
  • 适合蛋白功能研究者与模型开发者参考

深度学习在蛋白质功能与相互作用预测方面取得显著进展,但对蛋白功能机制的理解仍需更细致视角。为此,我们提出VenusX,首个面向残基、片段和结构域层面的细粒度功能注释与功能配对的大规模基准。该基准包含六类注释任务,涵盖残基级二分类、片段级多分类及配对功能相似性评分,用于识别关键活性位点、结合位点、保守位点、基序、结构域和表位。数据来自InterPro、BioLiP、SAbDab等开源数据库,总量超过87.8万条。通过设置三种序列相似性阈值的混合家族与跨家族划分,支持模型在分布内与分布外场景下的全面评估。我们测试了多种主流开源模型,包括预训练蛋白语言模型、序列-结构混合模型、基于结构的方法和比对技术,并在多个指标下报告性能,为后续研究提供坚实基础。代码与数据已公开于https://github.com/ai4protein/VenusX。

原文摘要 · Abstract (English)

Deep learning models have driven significant progress in predicting protein function and interactions at the protein level. While these advancements have been invaluable for many biological applications such as enzyme engineering and function annotation, a more detailed perspective is essential for understanding protein functional mechanisms and evaluating the biological knowledge captured by models. To address this demand, we introduce VenusX, the first large-scale benchmark for fine-grained functional annotation and function-based protein pairing at the residue, fragment, and domain levels. VenusX comprises three major task categories across six types of annotations, including residue-level binary classification, fragment-level multi-class classification, and pairwise functional similarity scoring for identifying critical active sites, binding sites, conserved sites, motifs, domains, and epitopes. The benchmark features over 878,000 samples curated from major open-source databases such as InterPro, BioLiP, and SAbDab. By providing mixed-family and cross-family splits at three sequence identity thresholds, our benchmark enables a comprehensive assessment of model performance on both in-distribution and out-of-distribution scenarios. For baseline evaluation, we assess a diverse set of popular and open-source models, including pre-trained protein language models, sequence-structure hybrids, structure-based methods, and alignment-based techniques. Their performance is reported across all benchmark datasets and evaluation settings using multiple metrics, offering a thorough comparison and a strong foundation for future research. Code and data are publicly available at https://github.com/ai4protein/VenusX.

蛋白质功能细粒度注释生物基准AI制药

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。