arXiv:2608.27459cs.AI2026-08

一个14B参数模型能回答41年《危险边缘》所有题目,且不依赖原始服务器。

Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model

论文配图:Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model
图 1 · 摘自论文原文
  • 用9GB的开源模型跑完整个Jeopardy!题库,实现本地化知识问答。
  • 答对67.0%的题目,事实类题正确率超85%,超越原版Watson表现。
  • 模型可随时间演化,不像当年的Watson被锁定在建造时刻。

2011年,IBM Watson凭借深度问答系统击败人类冠军,其知识来自运行在POWER7服务器集群上的百亿文档语料库,且无法迁移或复制。本文展示,如今类似的‘文化知识快照’已可便携且近乎免费获取。我们使用一个9 GB的开源模型(Qwen2.5-14B,4-bit)在完整开放的Jeopardy!题库上进行评估,该题库涵盖1984至2025年共41季的529,939道题目。据我们所知,这是首次在全量数据上运行模型。这些题目测试的是人类文化积累的通用知识,包括历史、语言、科学、文学与地理等,每题均有验证答案。模型在严格强制响应协议下以精确和模糊匹配方式答对67.0%的题目,事实类题目超过85%。我们认为训练数据暴露是所有系统共有的特性,而非语言模型独有缺陷;相反,Watson的语料仅针对过往题目构建,无法回答训练后的新问题。而在训练截止后播出的题目中,本地模型答对率65%,Claude Opus 4.8为95%,而Watson因构造原因得分为零。这表明能力可从服务器迁移到文件,且不被冻结于其生成时刻。

原文摘要 · Abstract (English)

In 2011, IBM's Watson was something like a sealed capsule of its era's queryable knowledge. Its DeepQA system defeated the strongest human Jeopardy! champions, but the knowledge that let it do so lived in a curated billion-document corpus running on a cluster of POWER7 servers, frozen at build time and impossible to move or copy. We show that the same kind of artifact, a snapshot of what a culture can answer, is now portable and essentially free. We evaluate a single 9 GB open-weight model (Qwen2.5-14B, 4-bit) against the complete open Jeopardy! clue dataset, 529,939 clues across all 41 broadcast seasons from 1984 to 2025. To our knowledge this is the first time a model has been run over the full corpus. The 41 years mark only how long the questions were collected. What they test is far older and broader: the accumulated body of human general knowledge a culture considers worth knowing, from ancient history and dead languages to science, literature, and geography, with a verified answer for every item. The model answers 67.0% of all clues under a strict forced-response protocol with exact and fuzzy matching, and exceeds 85% on factoid categories. We treat training-data exposure as something both systems share rather than a flaw unique to language models. Watson's case is in fact the more extreme one. Its corpus was assembled to contain Jeopardy answers and it was tuned on past clues, and it could not answer anything outside that curated distribution. The decisive test is whether a model can answer clues that did not exist when it was built. On clues aired after its training cutoff, the local model holds 65% and Claude Opus 4.8 holds 95%, while Watson by construction scores zero. The capability survives the move from a server room to a file you could seal in a time capsule, and unlike Watson it is not frozen to its own moment.

大模型知识问答开放数据可移植性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。