arXiv:2603.01464cs.AIcs.CL2026-03被引 1

用强化学习训练多轮蛋白搜索智能体,融合序列与文本信息提升医疗蛋白分析准确率

ProtRLSearch: A Multi-Round Multimodal Protein Search Agent with Large Language Models Trained via Reinforcement Learning

  • 基于多维奖励的强化学习,实现多轮交互式蛋白搜索
  • 在3000道多难度题目上达到82.7%准确率,显著优于单轮基线模型
  • 适合临床研究、疾病突变分析等需要融合序列与语义的场景

医疗场景中的蛋白质分析任务常需在序列约束下进行精准推理,涵盖致病突变功能解读、蛋白水平临床研究等。现有搜索代理多为单轮文本搜索,无法有效整合蛋白序列作为多模态输入。同时,依赖最终答案的强化学习监督导致搜索过程缺乏约束,关键词选择与推理路径偏差难以及时发现和纠正。为此,我们提出ProtRLSearch,一种通过多维奖励强化学习训练的多轮蛋白搜索代理,实时融合蛋白序列与文本信息生成高质量报告。为评估模型在真实蛋白查询中整合序列与文本多模态输入的能力,我们构建了ProtMCQs基准,包含3000道多选题,分为三个难度层级,涵盖从序列约束下的功能与表型变化推理,到结合多维序列特征、信号通路与调控网络的综合性蛋白推理任务。

原文摘要 · Abstract (English)

Protein analysis tasks arising in healthcare settings often require accurate reasoning under protein sequence constraints, involving tasks such as functional interpretation of disease-related variants, protein-level analysis for clinical research, and similar scenarios. To address such tasks, search agents are introduced to search protein-related information, providing support for disease-related variant analysis and protein function reasoning in protein-centric inference. However, such search agents are mostly limited to single-round, text-only modality search, which prevents the protein sequence modality from being incorporated as a multimodal input into the search decision-making process. Meanwhile, their reliance on reinforcement learning (RL) supervision that focuses solely on the final answer results in a lack of search process constraints, making deviations in keyword selection and reasoning directions difficult to identify and correct in a timely manner. To address these limitations, we propose ProtRLSearch, a multi-round protein search agent trained with multi-dimensional reward based RL, which jointly leverages protein sequence and text as multimodal inputs during real-time search to produce high quality reports. To evaluate the ability of models to integrate protein sequence information and text-based multimodal inputs in realistic protein query settings, we construct ProtMCQs, a benchmark of 3,000 multiple choice questions (MCQs) organized into three difficulty levels. The benchmark evaluates protein query tasks that range from sequence constrained reasoning about protein function and phenotype changes to comprehensive protein reasoning that integrates multi-dimensional sequence features with signal pathways and regulatory networks.

蛋白搜索强化学习多模态临床研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。