arXiv:2511.08866cs.CL2025-11被引 1

用自评估AI框架在生物医学文献中自动发现新假说

BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation

  • 设计可自评的LLM代理,分生成与评估两阶段迭代探索
  • 结合结构化与文本数据,使假说新颖性和相关性提升
  • 提供标准化基准,适合研究智能科研助手的学者

生物医学假说生成传统上依赖从海量科学文献中挖掘隐藏关联,常用方法如基于文献发现(LBD)。然而现有方法多依赖单一数据类型或预设提取模式,难以发现复杂新关联。大型语言模型(LLM)代理在信息检索、推理和生成方面展现出潜力,但其在生物医学假说生成中的应用受限于缺乏标准数据集与执行环境。为此,我们提出BioVerge综合基准与BioVerge Agent框架,构建面向前沿科学知识的标准化探索环境。数据集包含源自历史生物医学假说与PubMed文献的结构化与文本数据。BioVerge Agent采用基于ReAct的方法,含独立的生成与评估模块,实现假说提案的迭代生成与自评估。实验揭示:1)不同代理架构影响探索多样性与推理策略;2)结构化与文本信息源分别提供独特关键上下文;3)自评估显著提升假说的新颖性与相关性。

原文摘要 · Abstract (English)

Hypothesis generation in biomedical research has traditionally centered on uncovering hidden relationships within vast scientific literature, often using methods like Literature-Based Discovery (LBD). Despite progress, current approaches typically depend on single data types or predefined extraction patterns, which restricts the discovery of novel and complex connections. Recent advances in Large Language Model (LLM) agents show significant potential, with capabilities in information retrieval, reasoning, and generation. However, their application to biomedical hypothesis generation has been limited by the absence of standardized datasets and execution environments. To address this, we introduce BioVerge, a comprehensive benchmark, and BioVerge Agent, an LLM-based agent framework, to create a standardized environment for exploring biomedical hypothesis generation at the frontier of existing scientific knowledge. Our dataset includes structured and textual data derived from historical biomedical hypotheses and PubMed literature, organized to support exploration by LLM agents. BioVerge Agent utilizes a ReAct-based approach with distinct Generation and Evaluation modules that iteratively produce and self-assess hypothesis proposals. Through extensive experimentation, we uncover key insights: 1) different architectures of BioVerge Agent influence exploration diversity and reasoning strategies; 2) structured and textual information sources each provide unique, critical contexts that enhance hypothesis generation; and 3) self-evaluation significantly improves the novelty and relevance of proposed hypotheses.

假说生成LLM代理自评估生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。