arXiv:2504.12516cs.CL2025-04被引 578

测试智能体网页浏览能力的挑战性基准,聚焦信息搜寻的持久与创意。

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

  • 1266个问题需持续跨网页搜索复杂信息。
  • 答案简短可验证,评估重点在搜索策略而非回答长度。
  • 适合评估智能体在真实网络中持续探索的能力。

我们提出BrowseComp,一个简单但具有挑战性的基准,用于衡量智能体在网页浏览中的能力。该基准包含1,266个问题,要求智能体持续在网络中导航以获取难以找到且相互关联的信息。尽管问题难度高,但BrowseComp设计简洁,预测答案短小且易于与参考答案核对。虽然它未涵盖真实用户查询分布的所有挑战(如生成长答案或解决歧义),但有效评估了智能体在信息搜寻中展现的持久性与创造力这一核心能力。相关数据集和代码可在https://github.com/openai/simple-evals获取。

原文摘要 · Abstract (English)

We present BrowseComp, a simple yet challenging benchmark for measuring the ability for agents to browse the web. BrowseComp comprises 1,266 questions that require persistently navigating the internet in search of hard-to-find, entangled information. Despite the difficulty of the questions, BrowseComp is simple and easy-to-use, as predicted answers are short and easily verifiable against reference answers. BrowseComp for browsing agents can be seen as analogous to how programming competitions are an incomplete but useful benchmark for coding agents. While BrowseComp sidesteps challenges of a true user query distribution, like generating long answers or resolving ambiguity, it measures the important core capability of exercising persistence and creativity in finding information. BrowseComp can be found at https://github.com/openai/simple-evals.

网页浏览智能体评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。