arXiv:2608.22132cs.CLcs.AI2026-08中稿 · EMNLP

SSE-Bio通过动态检索与细粒度编辑提升生物医学多跳推理能力。

SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning

论文配图:SSE-Bio: A Structured Self-Evolving Agent with Agentic Retrieval Policy for Multi-Hop Biomedical Reasoning
图 1 · 摘自论文原文
  • 构建结构化状态,动态选择知识三元组与推理模板。
  • 在BioHopR上比最强基线高6.56分,多跳推理更准确。
  • 适合需要持续优化的生物医学问答系统开发者。

生物医学多跳问答需关联疾病、药物、蛋白质、表型等中间实体的证据。现有代理通常依赖静态检索流程或粗粒度提示重写,导致推理更新时指令漂移。我们提出SSE-Bio,一种具代理式检索策略的结构化自进化代理。它不全局重写指令,而是维护结构化状态,通过可训练代理策略选择性检索知识三元组和先前模板,并通过细粒度模板编辑优化推理记忆。为优化检索决策,引入基于组相对策略优化的代理训练策略,通过对比不同检索选择的决策组来改进代理。在三个生物医学多跳问答基准上的实验表明,SSE-Bio持续优于现有基线,在BioHopR上比最强自进化基线高出6.56绝对分数。

原文摘要 · Abstract (English)

Biomedical multi-hop question answering (QA) requires models to connect evidence across intermediate entities such as diseases, drugs, proteins, and phenotypes. Existing agents typically rely on static retrieval workflows or coarse-grained prompt rewriting, which can lead to instruction drift when reasoning procedures need to be updated. We propose SSE-Bio, a structured self-evolving agent with an agentic retrieval policy for multi-hop biomedical reasoning. Instead of globally rewriting agent instructions, SSE-Bio maintains a structured state, selectively retrieves knowledge triplets and prior templates through a trainable proxy policy, and improves its reasoning memory through fine-grained template editing. To optimise retrieval decisions, we introduce a proxy-training strategy based on group relative policy optimization, where the proxy is improved through decision-contrastive groups over alternative retrieval choices. Experiments on three biomedical multi-hop QA benchmarks show that SSE-Bio consistently outperforms existing baselines, achieving an improvement of 6.56 absolute points over the strongest self-evolving baseline on BioHopR.

多跳推理自进化生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。