首个面向韩语网络浏览任务的智能体评测基准,揭示大模型在韩语场景下的严重短板。
K-BrowseComp: A Web Browsing Agent Benchmark Grounded in Korean Contexts

- 构建400个韩语上下文网页浏览任务,300个经母语者验证
- 顶级大模型在韩语任务上准确率仅30%-45%,韩国自研模型不足10%
- 设计对抗性合成数据集,最强模型准确率仍仅26%,适合压力测试
前沿大模型评估正从基础能力(如指令遵循、推理)转向组合式、代理式能力,但韩语代理评测基准仍极为稀缺。我们提出K-BrowseComp,一个基于韩语场景的网页浏览智能体评测基准,包含400个问题。其中300个问题构成经过母语者手动构建与验证的K-BrowseComp-Verified子集。在该子集上,前沿大模型(包括GPT-5.5、DeepSeek-V4-Pro和GLM-5.1)准确率仅为30.00–45.67%,远低于原版BrowseComp表现;而韩国自主研发的专有大模型(来自韩国专属人工智能基础模型计划)准确率仅为0.00–10.33%。此外,我们利用难例少样本示例与失败模式靶向生成方法,构建了100个合成问题子集,以挖掘解决与生成任务之间的不对称性。在经过对抗过滤的合成诊断子集上,最强模型准确率仅为26.00%,我们单独报告此结果作为针对性压力测试。数据与代码已公开发布。
原文摘要 · Abstract (English)
Frontier model evaluations are shifting from foundational capabilities (e.g., instruction following and reasoning) toward compositional, agentic ones, but Korean agentic benchmarks remain scarce. We introduce K-BrowseComp, a web-browsing agent benchmark grounded in Korean contexts, consisting of 400 problems. The 300-problem K-BrowseComp-Verified subset is manually constructed and validated by native Korean speakers. On this subset, frontier LLMs, including GPT-5.5, DeepSeek-V4-Pro, and GLM-5.1, reach only 30.00--45.67\%, a substantial drop from BrowseComp, while Korean LLMs released through Korea's Proprietary AI Foundation Model program obtain only 0.00--10.33\%. We further construct a 100-problem synthetic split using hard few-shot exemplars and failure-mode-targeted generation to exploit the asymmetry between solving and creating web browsing problems. On the adversarially filtered synthetic diagnostic split, the strongest model reaches only 26.00\%, and we report this split separately as a targeted stress test. We publicly release our data and code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。