用侦探游戏评测大模型推理,发现其逐步逼近人类水平。
Watson & Holmes: A Naturalistic Benchmark for Comparing Human and LLM Reasoning
- 用叙事性推理任务模拟真实思考过程,支持开放问答。
- 2025年9个月间模型表现从人类后四分之一升至前5%。
- 长案件中模型退化,初期证据少时推理模型占优。
现有AI推理评测难以揭示其与人类推理在自然情境中的相似性。我们改编了‘Watson & Holmes’侦探桌游,构建了一个新基准:通过逐步呈现叙事证据、开放性问题和自由语言回答,评估推理能力。开发了自动化评分系统,并经由人工评估验证,实现可扩展、可复现的性能评估。结果显示,模型性能随时间显著提升:2025年九个月内,模型表现从人类比较组的下四分之一跃升至约前5%。其中约一半进步源于各代模型的持续迭代,另一半则归因于面向推理优化的模型架构带来的显著跃升。系统性差异主要出现在特定谜题特征上:模型在较长案件(1900-4000词)中表现下降;而在案件初期证据稀缺时,推理导向模型在归纳推理上具优势。
原文摘要 · Abstract (English)
Existing benchmarks for AI reasoning provide limited insight into how closely these capabilities resemble human reasoning in naturalistic contexts. We present an adaptation of the Watson & Holmes detective tabletop game as a new benchmark designed to evaluate reasoning performance using incrementally presented narrative evidence, open-ended questions and unconstrained language responses. An automated grading system was developed and validated against human assessors to enable scalable and replicable performance evaluation. Results show a clear improvement in AI model performance over time. Over nine months of 2025, model performance rose from the lower quartile of the human comparison group to approximately the top 5%. Around half of this improvement reflects steady advancement across successive model releases, while the remainder corresponds to a marked step change associated with reasoning-oriented model architectures. Systematic differences in the performance of AI models compared to humans, dependent on features of the specific detection puzzle, were mostly absent with the exception of a fall in performance for models when solving longer cases (case lengths being in the range of 1900-4000 words), and an advantage at inductive reasoning for reasoning models at early stages of case solving when evidence was scant.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。