让蛋白质模型学会自我纠错,提升生物序列理解能力。
Reflection Pretraining Enables Token-Level Self-Correction in Biological Sequence Models
- 引入反思预训练,生成额外思考标记实现中间推理。
- 在蛋白质语言模型上实现自纠错,准确率显著提升。
- 适合生物序列分析、蛋白设计等研究者使用。
链式思维(CoT)提示已显著提升大语言模型在自然语言处理中的任务解决能力。与标准提示不同,CoT鼓励模型生成中间推理步骤,即非答案标记,以引导模型得出更准确的最终输出。这些中间步骤支持复杂推理过程,如错误纠正、记忆管理、未来规划和自我反思。然而,将CoT应用于非自然语言领域(如蛋白质和RNA语言模型)仍不可行,主要由于其标记空间表达力有限(如氨基酸标记)。本文提出并定义了语言表达力概念:给定语言使用其标记和语法编码信息的能力。我们证明蛋白质语言的表达力受限严重制约了CoT式推理的应用。为此,我们首次在生物序列模型中引入反思预训练,使模型通过生成超出简单答案标记的辅助“思考标记”进行中间推理。理论上,我们证明扩充标记集显著增强生物语言表达力,从而提升模型整体推理能力。实验上,该预训练方法使蛋白质模型具备自纠错能力,并带来显著性能提升。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) prompting has significantly advanced task-solving capabilities in natural language processing with large language models. Unlike standard prompting, CoT encourages the model to generate intermediate reasoning steps, non-answer tokens, that help guide the model toward more accurate final outputs. These intermediate steps enable more complex reasoning processes such as error correction, memory management, future planning, and self-reflection. However, applying CoT to non-natural language domains, such as protein and RNA language models, is not yet possible, primarily due to the limited expressiveness of their token spaces (e.g., amino acid tokens). In this work, we propose and define the concept of language expressiveness: the ability of a given language, using its tokens and grammar, to encode information. We show that the limited expressiveness of protein language severely restricts the applicability of CoT-style reasoning. To overcome this, we introduce reflection pretraining, for the first time in a biological sequence model, which enables the model to engage in intermediate reasoning through the generation of auxiliary "thinking tokens" beyond simple answer tokens. Theoretically, we demonstrate that our augmented token set significantly enhances biological language expressiveness, thereby improving the overall reasoning capacity of the model. Experimentally, our pretraining approach teaches protein models to self-correct and leads to substantial performance gains compared to standard pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。