arXiv:2605.06177cs.AI2026-05

BioMedArena统一生物医学研究智能体评估环境,解决复现难题。

BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

论文配图:BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents
图 1 · 摘自论文原文
  • 拆分六层评估流程,支持快速集成新模型与工具
  • 166个基准+75个工具,最佳配置提升15.01个百分点
  • 开源框架,适合医疗AI研究者快速对比与部署

当前深度研究智能体的复现与比较困难:相同主干模型在相同基准上报告的准确率因评估工具链差异而不同,且集成新模型需数周定制工程。为此,我们提出BioMedArena——一个开源工具包,旨在解决深度研究智能体领域的可复现性问题。该工具包解耦了生物医学智能体评估的六个层次:基准加载、工具暴露、工具选择、评估模式、上下文管理与评分,并提供了166个生物医学基准和75个跨9类功能的生物医学工具。新增模型、基准或工具仅需编写几行适配代码。除评估基础设施外,还包含高质量参考组件:6种智能体评估框架(含提出的Mutual-Evolve)与6种上下文管理策略,可任意搭配于任一主干模型。使用这些组件后,所有12个主干模型性能显著提升;在8个代表性生物医学基准上,最优配置的主干模型平均超越先前最先进水平15.01个百分点。工具包、配置及任务追踪记录已公开于https://github.com/AI-in-Health/BioMedArena。

原文摘要 · Abstract (English)

Reproducing and comparing deep research agents today is hard: the same backbone evaluated on the same benchmark can report different accuracies across papers because the harness and tool registry differ, and integrating a new model into a comparable evaluation surface costs weeks of model-specific engineering. These are symptoms of a broader reproducibility problem in deep research agent research. Here, we introduce BioMedArena, an open-source toolkit that addresses this reproducibility gap and provides an arena for comparing deep research agents under a shared evaluation environment. BioMedArena decouples six layers of biomedical agent evaluation -- benchmark loading, tool exposure, tool selection, harness mode, context management, and scoring -- and exposes 166 biomedical benchmarks and 75 biomedical tools across 9 functional families. Adding a new model, benchmark, or tool can be accomplished with a few-line provider adapter. Beyond evaluation infrastructure, BioMedArena ships a library of high-quality reference components: 6 agent harnesses (including our proposed Mutual-Evolve) and 6 context-management strategies, any of which can be equipped on any backbone. Equipping these components substantially improves all 12 backbones; on each of 8 representative biomedical benchmarks, the best equipped backbone surpasses prior state-of-the-art (SOTA), by 15.01 percentage points on average. The toolkit, configurations, and per-task traces are available at https://github.com/AI-in-Health/BioMedArena.

智能体生物医学评估框架开源工具

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。