评测大模型在中文网页浏览中的推理与检索能力,发现多数模型表现极差。
BrowseComp-ZH: Benchmarking Web Browsing Ability of Large Language Models in Chinese
- 构建289个跨领域多跳问题,答案可验证且唯一。
- 超20个顶尖模型测试,最高准确率仅42.9%。
- 适合研究中文信息获取、智能代理的学者参考。
随着大语言模型(LLMs)演变为工具使用代理,实时网页浏览能力已成为衡量其推理与检索能力的关键标准。现有基准如BrowseComp主要面向英文环境,忽视了中文信息生态在语言、基础设施及审查机制上的复杂性。为填补这一空白,我们提出BrowseComp-ZH,一个专为评估中文网络环境下大模型代理能力而设计的高难度基准。该基准包含289个涵盖11个不同领域的多跳问题,每个问题均基于简短、客观且可验证的答案(如日期、数字或专有名词)逆向构建。采用两阶段质量控制流程以确保问题难度和答案唯一性。我们在该基准上对超过20个前沿语言模型和智能搜索系统进行了评测。尽管这些模型具备强大的对话与检索能力,但多数表现不佳:大量模型准确率低于10%,仅有少数超过20%。即使最佳系统OpenAI的DeepResearch也仅达到42.9%的准确率。结果表明BrowseComp-ZH难度极高,成功需兼具高效检索策略、复杂推理与信息整合能力,当前模型仍难以掌握。我们的数据集、构建指南与评测结果已公开于https://github.com/PALIN2018/BrowseComp-ZH。
原文摘要 · Abstract (English)
As large language models (LLMs) evolve into tool-using agents, the ability to browse the web in real-time has become a critical yardstick for measuring their reasoning and retrieval competence. Existing benchmarks such as BrowseComp concentrate on English and overlook the linguistic, infrastructural, and censorship-related complexities of other major information ecosystems -- most notably Chinese. To address this gap, we introduce BrowseComp-ZH, a high-difficulty benchmark purpose-built to comprehensively evaluate LLM agents on the Chinese web. BrowseComp-ZH consists of 289 multi-hop questions spanning 11 diverse domains. Each question is reverse-engineered from a short, objective, and easily verifiable answer (e.g., a date, number, or proper noun). A two-stage quality control protocol is applied to strive for high question difficulty and answer uniqueness. We benchmark over 20 state-of-the-art language models and agentic search systems on our proposed BrowseComp-ZH. Despite their strong conversational and retrieval capabilities, most models struggle severely: a large number achieve accuracy rates below 10%, and only a handful exceed 20%. Even the best-performing system, OpenAI's DeepResearch, reaches just 42.9%. These results demonstrate the considerable difficulty of BrowseComp-ZH, where success demands not only effective retrieval strategies, but also sophisticated reasoning and information reconciliation -- capabilities that current models still struggle to master. Our dataset, construction guidelines, and benchmark results have been publicly released at https://github.com/PALIN2018/BrowseComp-ZH.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。