arXiv:2606.21959cs.CL2026-06

构建首个面向未解医学问题的智能体评测基准,揭示模型在关键推理中工具失效的问题。

OpenBioRQ: Unsolved Biomedical Research Questions for Agents

  • 设计12,553个跨领域的未解医学问题,以真实证据验证开放性答案
  • 实测顶尖模型在最难问题上仅能解决17%~60%,表明任务极具挑战性
  • 发现模型在高难度问题中放弃使用工具,暴露智能体系统性失效

当前智能体模型虽极少虚构引用(超过99%链接可解析),但约15.9%存在引文错误。现有评测忽略此缺陷:当问题有固定答案时,模型可直接复现对应文献而非验证其支持性。本文提出 extbf{ exttt{OpenBioRQ}},一个包含12,553个跨12个领域的未解生物医学问题的检索增强型智能体基准,将开放问题作为忠实性与不回答的探测器。该基准首次结合智能体多步操作设置与无答案键的开放问题,通过真实后续证据验证开放性,而非依赖模型参数知识。难度基于三个开源参考模型无法解答的问题确定,非主观标注。在最困难子集上,同源模型仅能解决约17%,而三个独立前沿代理(Gemini-3-Pro、Opus-4.7、GPT-5.5)得分范围为29%-60%。该基准具有高难度、非饱和性(最佳模型仍有33%-40%未解)和能力区分力。此外,观察到在最难问题上出现智能体崩溃现象:模型停止调用工具。对最易崩溃模型,禁用工具后成绩几乎不变——工具在最需使用时反而失效。引入每题检查清单后,评审者间一致性从斯皮尔曼相关0.35提升至0.82。

原文摘要 · Abstract (English)

A working citation looks like proof -- but the fact that a link resolves does not mean the cited paper supports the claim. I find that current agentic models rarely fabricate citations (over $99\%$ resolve), yet roughly $15.9\%$ link to the wrong paper. Existing benchmarks miss this failure mode: when a question has a fixed answer key, a model can reproduce the expected source from that key rather than independently verifying that the source supports the claim. I introduce \textbf{\openbiorq{}}, a retrieval-grounded agentic benchmark of $12{,}553$ unsolved biomedical research questions across $12$ domains that treats open questions as a faithfulness-and-abstention probe. To my knowledge, this is the first biomedical benchmark to combine an agentic setting -- where the model must issue multiple tool calls -- with unsolved questions that have no answer key. Openness is verified against real follow-up evidence rather than a model's parametric knowledge. Difficulty is empirical: I anchor it on questions that three open-weight reference models fail to answer, rather than on subjective hardness labels. On this hardest subset, held-out models from the same lineage as the difficulty anchors solve only ~17%, while three independent frontier agents (Gemini-3-Pro, Opus-4.7, GPT-5.5) span a wide 29-60% range. The benchmark is thus hard, non-saturating (the best agent still leaves ~33-40\% unsolved), and discriminating across capability tiers. Beyond difficulty, I observe agentic collapse on the hardest questions, where agents stop using their tools. For the most collapse-prone model, blocking tool access entirely barely changes its score -- so tools stop paying off exactly where they are needed most. A frozen per-question checklist raises inter-judge agreement from Spearman 0.35 to 0.82.

智能体评测医学问答工具使用真实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。