arXiv:2509.26619cs.CLcs.AI2025-09被引 1

自动上网搜难例,低成本构建真实挑战基准

Searching the Internet for Challenging Benchmarks at Scale

  • 把互联网当话题池,用强化学习策略高效找最难题目
  • 只查6%内容就找到难点,成本比全量检查低100倍
  • 适合评估模型真实短板,尤其对多语言任务有用

许多静态基准正面临饱和:随着模型快速进步,它们在固定测试集上几乎达到完美分数,难以暴露真实弱点——甚至专家设计的挑战集也会在快速优化后失效。我们提出一个全自动框架,可大规模搜索互联网以构建无需人工标注的挑战性基准。核心思路是将互联网视为庞大话题空间,并将搜索建模为多臂老虎机问题,每个话题的难度仅通过昂贵的采样与评估才能揭示。采用epsilon-greedy策略,在仅探索6%搜索空间的情况下,成功识别出最具挑战性的主题,相比全面评估节省了100倍成本。我们在机器翻译和知识问答任务上验证,发现所发现的难度在独立指标(GEMBA-SQA 和 MetricX)、多种语言及不同模型间均具鲁棒性。

原文摘要 · Abstract (English)

Many static benchmarks are beginning to saturate: as models rapidly improve, they achieve near-perfect scores on fixed test sets, leaving little headroom to expose genuine model weaknesses -- and even expert-curated challenge sets quickly saturate after hillclimbing. We present a fully automatic framework that searches the Internet at scale to construct challenging benchmarks without human curation. The key insight is to model the Internet as a vast space of topics and formalize the search as a multi-armed bandit problem, where each topic's difficulty is revealed only through expensive sample-and-evaluate queries. Our epsilon-greedy strategy identifies the most challenging topics while exploring only 6% of the search space -- a 100 times cost reduction over exhaustive evaluation. We validate on machine translation and knowledge question answering, confirming that discovered difficulty is robust across independent metrics (GEMBA-SQA and MetricX), languages, and models.

基准测试自动构建多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。