arXiv:2510.23127cs.AI2025-10

用结构化生物信息学上下文替代原始序列,显著提升科学大模型的生物推理能力。

Lost in Tokenization: Context as the Key to Unlocking Biomolecular Understanding in Scientific LLMs

  • 用专家工具生成的结构化上下文替代原始序列输入
  • 仅使用上下文时模型性能超越序列或组合输入
  • 适合希望提升生物序列推理效率的研究者

科学大语言模型(Sci-LLMs)在加速生物学发现方面展现出巨大潜力,但在处理原始生物分子序列时面临分词困境:将序列视为特殊语言可能丢失功能基序信息,而作为独立模态则引入对齐难题。本文挑战以序列为中心的范式,提出通过提供来自成熟生物信息学工具的高层结构化上下文,绕过对低层噪声序列的直接解析。我们在多个生物推理任务中系统比较了主流Sci-LLMs在三种输入模式下的表现:仅序列、仅上下文、以及两者结合。结果表明,仅使用上下文的模式始终显著优于其他方式;甚至在加入原始序列后,性能反而下降,说明原始序列会带来信息干扰,即便对具备专用分词方案的模型也是如此。这表明现有Sci-LLMs的核心优势不在于从头理解分子语法,而在于对结构化人类可读知识的深层推理。因此,我们主张将Sci-LLMs重新定位为基于专家知识的强大推理引擎。本研究为新型混合科学智能体奠定基础,推动研发重点从直接序列解读转向高层知识融合。代码已开源:https://github.com/opendatalab-raiser/CoKE。

原文摘要 · Abstract (English)

Scientific Large Language Models (Sci-LLMs) have emerged as a promising frontier for accelerating biological discovery. However, these models face a fundamental challenge when processing raw biomolecular sequences: the tokenization dilemma. Whether treating sequences as a specialized language, risking the loss of functional motif information, or as a separate modality, introducing formidable alignment challenges, current strategies fundamentally limit their reasoning capacity. We challenge this sequence-centric paradigm by positing that a more effective strategy is to provide Sci-LLMs with high-level structured context derived from established bioinformatics tools, thereby bypassing the need to interpret low-level noisy sequence data directly. Through a systematic comparison of leading Sci-LLMs on biological reasoning tasks, we tested three input modes: sequence-only, context-only, and a combination of both. Our findings are striking: the context-only approach consistently and substantially outperforms all other modes. Even more revealing, the inclusion of the raw sequence alongside its high-level context consistently degrades performance, indicating that raw sequences act as informational noise, even for models with specialized tokenization schemes. These results suggest that the primary strength of existing Sci-LLMs lies not in their nascent ability to interpret biomolecular syntax from scratch, but in their profound capacity for reasoning over structured, human-readable knowledge. Therefore, we argue for reframing Sci-LLMs not as sequence decoders, but as powerful reasoning engines over expert knowledge. This work lays the foundation for a new class of hybrid scientific AI agents, repositioning the developmental focus from direct sequence interpretation towards high-level knowledge synthesis. The code is available at https://github.com/opendatalab-raiser/CoKE.

科学大模型生物信息学知识推理上下文增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。