arXiv:2507.00953q-bio.BMcs.AI2025-07被引 4

将生物序列视为有结构语义的‘语言’,重新思考大模型在蛋白质等系统中的应用。

From Sentences to Sequences: Rethinking Languages in Biological System

  • 以生物分子三维结构作为语义核心,重构生物语言建模思路
  • 验证自回归生成在蛋白质序列建模中的有效性
  • 适合对生物序列与结构关系感兴趣的跨领域研究者

大语言模型在自然语言处理中的范式已成功应用于蛋白质、RNA和DNA等生物语言建模。尽管自回归生成方式与评估指标被直接迁移,但自然语言与生物语言在内在结构相关性上存在根本差异。为此,本文重新审视生物系统中的语言概念,将生物分子的三维结构视为句子的语义内容,并考虑残基或碱基间的强相关性,强调结构评估的重要性,同时证明自回归范式在生物语言建模中的适用性。代码可在 github.com/zjuKeLiu/RiFold 获取。

原文摘要 · Abstract (English)

The paradigm of large language models in natural language processing (NLP) has also shown promise in modeling biological languages, including proteins, RNA, and DNA. Both the auto-regressive generation paradigm and evaluation metrics have been transferred from NLP to biological sequence modeling. However, the intrinsic structural correlations in natural and biological languages differ fundamentally. Therefore, we revisit the notion of language in biological systems to better understand how NLP successes can be effectively translated to biological domains. By treating the 3D structure of biomolecules as the semantic content of a sentence and accounting for the strong correlations between residues or bases, we highlight the importance of structural evaluation and demonstrate the applicability of the auto-regressive paradigm in biological language modeling. Code can be found at \href{https://github.com/zjuKeLiu/RiFold}{github.com/zjuKeLiu/RiFold}

生物语言结构建模自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。