arXiv:2508.17580cs.CLcs.AI2025-08被引 11

用未解问题评估大模型,让测试更真实且有实际价值。

UQ: Assessing Language Models on Unsolved Questions

  • 以未解决的现实问题构建动态评测集,突破传统基准的难度-真实性矛盾。
  • 顶尖模型仅在15%的问题上通过验证,体现真实挑战性。
  • 适合关注模型真实推理能力与知识边界的研究者使用。

基准测试推动人工智能研究进展。一个有效的基准应兼具难度与现实意义:问题需挑战前沿模型,同时反映真实应用场景。然而当前范式面临难度-真实性矛盾:考试式基准常人为设难,缺乏现实价值;基于真实用户交互的基准则多集中于简单高频问题。本文提出全新范式:在未解问题上评估模型。不采用一次性静态评测,而是持续异步评估,结合验证者筛选与社区验证。我们推出UQ,包含500个来自Stack Exchange的多样化难题,涵盖计算机理论、数学、科幻、历史等领域,考察推理、事实性与信息检索能力。这些未解问题天然具有难度和现实背景,解决它们可直接产生实际价值。贡献包括:(1)UQ数据集及采集流程,融合规则过滤、大模型判断与人工审查,确保问题质量(如定义清晰、难度高);(2)UQ验证者机制,利用生成器-验证者差距提供评估信号,并预筛候选答案供人工审核;(3)开放平台UQ-Platform,支持专家集体验证问题与答案。顶级模型仅在15%的问题上通过验证,初步人工验证已发现其中正确答案。UQ为评估前沿模型在真实开放问题上的表现开辟路径,成功将推动人类知识边界。项目地址:https://uq.stanford.edu。

原文摘要 · Abstract (English)

Benchmarks shape progress in AI research. A useful benchmark should be both difficult and realistic: questions should challenge frontier models while also reflecting real-world usage. Yet, current paradigms face a difficulty-realism tension: exam-style benchmarks are often made artificially difficult with limited real-world value, while benchmarks based on real user interaction often skew toward easy, high-frequency problems. In this work, we explore a radically different paradigm: assessing models on unsolved questions. Rather than a static benchmark scored once, we curate unsolved questions and evaluate models asynchronously over time with validator-assisted screening and community verification. We introduce UQ, a testbed of 500 challenging, diverse questions sourced from Stack Exchange, spanning topics from CS theory and math to sci-fi and history, probing capabilities including reasoning, factuality, and browsing. UQ is difficult and realistic by construction: unsolved questions are often hard and naturally arise when humans seek answers, thus solving them yields direct real-world value. Our contributions are threefold: (1) UQ-Dataset and its collection pipeline combining rule-based filters, LLM judges, and human review to ensure question quality (e.g., well-defined and difficult); (2) UQ-Validators, compound validation strategies that leverage the generator-validator gap to provide evaluation signals and pre-screen candidate solutions for human review; and (3) UQ-Platform, an open platform where experts collectively verify questions and solutions. The top model passes UQ-validation on only 15% of questions, and preliminary human verification has already identified correct answers among those that passed. UQ charts a path for evaluating frontier models on real-world, open-ended challenges, where success pushes the frontier of human knowledge. We release UQ at https://uq.stanford.edu.

大模型评测未解问题真实场景开放问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。