arXiv:2503.00089q-bio.QMcs.AI2025-03ICML被引 18

提出新方法提升蛋白质结构分词效率,性能比ESM3平均高6.31%

Protein Structure Tokenization: Benchmarking and New Recipe

  • 设计细粒度局部结构评估框架,突破传统全局结构评价局限
  • AminoAseed策略使代码本利用率提升124%,敏感度增12.83%
  • 适合蛋白质结构建模、多模态模型融合的研究者参考

近年来蛋白质结构分词方法快速发展,将三维结构转化为离散或连续表示,使语言建模等技术可直接应用于蛋白质结构,并支持多模态模型整合序列与功能文本。然而,由于缺乏统一评估框架,现有方法的能力与局限尚不清晰。本文提出StructTokenBench,一个聚焦细粒度局部子结构而非全局结构的全面评估框架。评估发现无单一模型在所有维度占优。观察到代码本使用不足后,提出AminoAseed策略,通过优化代码本梯度更新,合理平衡代码本大小与维度,显著提升利用效率与质量。相比领先模型ESM3,该方法在24个监督任务上平均性能提升6.31%,敏感度与利用率分别提高12.83%和124.03%。源码与模型权重已公开于https://github.com/KatarinaYuan/StructTokenBench。

原文摘要 · Abstract (English)

Recent years have witnessed a surge in the development of protein structural tokenization methods, which chunk protein 3D structures into discrete or continuous representations. Structure tokenization enables the direct application of powerful techniques like language modeling for protein structures, and large multimodal models to integrate structures with protein sequences and functional texts. Despite the progress, the capabilities and limitations of these methods remain poorly understood due to the lack of a unified evaluation framework. We first introduce StructTokenBench, a framework that comprehensively evaluates the quality and efficiency of structure tokenizers, focusing on fine-grained local substructures rather than global structures, as typical in existing benchmarks. Our evaluations reveal that no single model dominates all benchmarking perspectives. Observations of codebook under-utilization led us to develop AminoAseed, a simple yet effective strategy that enhances codebook gradient updates and optimally balances codebook size and dimension for improved tokenizer utilization and quality. Compared to the leading model ESM3, our method achieves an average of 6.31% performance improvement across 24 supervised tasks, with sensitivity and utilization rates increased by 12.83% and 124.03%, respectively. Source code and model weights are available at https://github.com/KatarinaYuan/StructTokenBench

蛋白质结构分词方法多模态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。