arXiv:2505.20354cs.CLcs.AI2025-05EMNLP被引 7

提出检索增强方法,解决蛋白文本模型评估偏差问题。

Rethinking Text-based Protein Understanding: Retrieval or LLM?

  • 用生物实体重构数据集,避免评测中的数据泄露。
  • 检索增强法在蛋白到文本生成上优于微调大模型。
  • 无需训练即可高效准确生成,适合快速应用开发。

近年来,蛋白-文本模型因其在蛋白生成与理解方面的潜力而受到广泛关注。现有方法通过持续预训练和多模态对齐将蛋白知识融入大语言模型,实现对文本描述与蛋白序列的同步理解。通过对现有模型架构及文本基蛋白理解基准的深入分析,我们发现当前基准存在显著的数据泄露问题。此外,源自自然语言处理的传统指标无法准确评估模型在此领域的性能。为解决这些局限,我们重新组织了现有数据集,并基于生物实体提出了新的评估框架。受此启发,我们提出一种检索增强方法,在蛋白到文本生成任务中显著优于微调的大语言模型,且在无训练场景下表现出高精度与高效率。代码与数据可在 https://github.com/IDEA-XL/RAPM 获取。

原文摘要 · Abstract (English)

In recent years, protein-text models have gained significant attention for their potential in protein generation and understanding. Current approaches focus on integrating protein-related knowledge into large language models through continued pretraining and multi-modal alignment, enabling simultaneous comprehension of textual descriptions and protein sequences. Through a thorough analysis of existing model architectures and text-based protein understanding benchmarks, we identify significant data leakage issues present in current benchmarks. Moreover, conventional metrics derived from natural language processing fail to accurately assess the model's performance in this domain. To address these limitations, we reorganize existing datasets and introduce a novel evaluation framework based on biological entities. Motivated by our observation, we propose a retrieval-enhanced method, which significantly outperforms fine-tuned LLMs for protein-to-text generation and shows accuracy and efficiency in training-free scenarios. Our code and data can be seen at https://github.com/IDEA-XL/RAPM.

蛋白生成检索增强评估偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。