用维多利亚时代文本训练模型,让幻觉成为理解历史的新方法
TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

- 用1800-1875年维多利亚时期文本训练12亿参数模型,实现时间隔离
- 在维多利亚散文上困惑度降低45.4%,优于当代大模型的表面表现
- 模型生成的历史类比可揭示19世纪认知框架,适合人文研究者使用
大型语言模型因训练数据集中于当代文本而存在时间过曝问题,难以准确叙述过去。本文提出TimeCapsule,一个1.2B参数的类LLaMA因果模型,仅用维多利亚时期(1800–1875)文本训练,作为知识上孤立的生成档案。定量评估显示,其在保留的维多利亚散文上的困惑度比GPT-2基线降低45.4%;尽管更大规模的现代因果模型在原始困惑度上更低,但缺乏时间隔离性。TimeCapsule展现出计算意义上的意义建构能力,能对陌生现代概念生成历史合理类比(如将计算机描述为“过度发达的肺”)。两位人文学学者进行定性诠释探查时,约40%的真实维多利亚文本被误判为机器生成,暴露出真实性危机。我们主张,对未来的结构性无知使幻觉转化为对十九世纪本体论的解释性探针。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliable narrators of the past. We present TimeCapsule, a 1.2B-parameter LLaMA-style causal model trained exclusively on Victorian texts (1800-1875) as an epistemologically isolated generative archive. Quantitative evaluation shows a 45.4% perplexity reduction over a GPT-2 baseline on held-out Victorian prose, while larger contemporary causal models achieve lower raw perplexity through broader pretraining but lack temporal isolation. TimeCapsule exhibits computational sensemaking, generating historically plausible analogical explanations for unfamiliar modern concepts (e.g., describing a computer as a "hypertrophied lung"). A qualitative hermeneutic probe with two humanities scholars revealed a crisis of authenticity, as both misclassified approximately 40% of genuine Victorian excerpts as machine-produced. We argue that structural ignorance of the future transforms hallucinations into interpretive probes of nineteenth-century ontologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。