学生自建评测集,检验AI在人文学科中的真实表现。
Teaching AI Through Benchmark Construction: QuestBench as a Course-Based Practice for Accountable Knowledge Work

- 让学生将专业知识转化为可验证的高质量问题,构建评测任务。
- 13个系统平均通过率仅16.85%,顶尖模型GPT-5.5也仅达57.58%。
- 适合教育者和研究者用于培养批判性AI使用能力。
随着AI融入日常学习,多数课程仅将其作为生产力工具教学。本文提出,应通过评测集构建训练学生评估AI、理解自身在知识判断中的角色。我们设计了一门基于课程的实践:学生将学科知识转化为可验证的专家级问题,互评设计质量,并评估深度研究系统在这些任务上的表现。该实践生成了名为QuestBench的评测集,包含跨14个人文与社会科学领域的256个问题。在该评测上,13个系统平均问题通过率仅为16.85%,表现最佳的GPT-5.5达到57.58%。失败案例揭示了流畅且有来源支持的回答仍可能遗漏关键查询、来源或证据标准。五名学生反思表明,该过程帮助他们将专业知识视为评判AI输出的基础。本文提供可复用的教学场景与数据集(链接:https://huggingface.co/datasets/PKUAIWeb/QuestBench/tree/main),探讨如何在AI融入学习与工作时保持学生的责任主体地位。
原文摘要 · Abstract (English)
As AI becomes part of everyday learning, many courses teach students to use it mainly as a productivity tool: how to prompt, search, summarize, write, code, and use tools more efficiently. We argue that AI education also needs a setting in which students learn to test AI and understand their own role in judging machine-produced knowledge. To this end, we introduce a course-based practice that teaches AI through benchmark construction, using deep research systems as a concrete example of AI-era knowledge work. Students turn disciplinary knowledge into verifiable expert-level questions, review one another's designs for ambiguity and shortcuts, and evaluate AI systems on the resulting tasks. This activity gives students direct exposure to a powerful tool while asking them to specify what a trustworthy answer would require. The produced benchmark, QuestBench, consists of 256 questions across 14 humanities and social-science domains. Evaluation on QuestBench shows that student-designed tasks reveal hidden failures in current deep research systems: across thirteen evaluated systems, the mean question-level pass rate is only 16.85%, and the best-performing system, GPT-5.5, reaches a 57.58% pass rate. The failures are educationally useful because they show how fluent, source-backed answers can still miss the right query, source, term, or evidence standard. Reflections from five student contributors suggest that benchmark construction can help students see professional knowledge not only as content AI may retrieve, but as the basis for judging AI outputs. We present QuestBench as a benchmark artifact and as a reusable classroom setting for a larger educational question: how students can remain responsible knowledge actors as AI enters learning and professional work. The dataset is available at https://huggingface.co/datasets/PKUAIWeb/QuestBench/tree/main.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。