arXiv:2601.03604cs.AI2026-01ACL被引 1

用工具调用增强蛋白质功能预测,效果比纯文本推理提升103%。

Interleaved Tool-Call Reasoning for Protein Function Understanding

  • 引入领域工具生成可验证的中间证据,替代长篇文本推理
  • 在4个基准上平均性能提升103%,显著优于纯文本模型
  • 适合需要可靠生物知识推理的研究者和药物研发人员

大语言模型在数学与编程等符号领域中展现出链式思维推理的有效性。然而,本研究发现,直接将此类文本推理范式迁移至蛋白质功能理解效果不佳:强化学习主要放大表面关键词模式,无法引入新生物学知识,导致泛化能力有限。我们认为,蛋白质功能预测是依赖外部生物先验和计算工具的知识密集型科学任务,而非纯粹内部推理。为此,我们提出PFUA——一种工具增强的蛋白质推理代理,统一问题分解、工具调用与有根据的答案生成。不同于依赖长而无约束的推理过程,PFUA通过整合领域特定工具产生可验证的中间证据。在四个基准上的实验表明,PFUA始终优于仅使用文本的推理模型,平均性能提升103%。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have highlighted the effectiveness of chain-of-thought reasoning in symbolic domains such as mathematics and programming. However, our study shows that directly transferring such text-based reasoning paradigms to protein function understanding is ineffective: reinforcement learning mainly amplifies superficial keyword patterns while failing to introduce new biological knowledge, resulting in limited generalization. We argue that protein function prediction is a knowledge-intensive scientific task that fundamentally relies on external biological priors and computational tools rather than purely internal reasoning. To address this gap, we propose PFUA, a tool-augmented protein reasoning agent that unifies problem decomposition, tool invocation, and grounded answer generation. Instead of relying on long unconstrained reasoning traces, PFUA integrates domain-specific tools to produce verifiable intermediate evidence. Experiments on four benchmarks demonstrate that PFUA consistently outperforms text-only reasoning models with an average performance improvement of 103%.

蛋白质预测工具调用推理增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。