72亿参数历史语言模型,专攻1913年前英文文本理解。
Pretraining Language Models on Historical Text

- 用540亿词元历史语料训练,严防时间泄露。
- 构建两个指令微调数据集,确保回答扎根史料。
- 推出新评测基准,验证模型历史一致性与准确性。
我们提出TypewriterLM,一个7.24B参数的历史语言模型,仅在1913年前的英文文本上预训练。构建历史语言模型面临数据质量、时间泄露、后训练流程一致性及评估可靠性等挑战。为此,我们构建了540亿词元的TypewriterCorpus,来自多样档案与语言标注资源,经过严格清洗和泄漏防护。我们提出基于词汇的指令微调框架,使模型输出直接基于历史文献。据此构建了两个历史指令微调数据集:History-LIMA与History-SelfInstruct。为评估能力与时间一致性,我们设计History-Event基准套件,用于检测任务完成度、时间定位准确性和数据泄露。我们已开源TypewriterLM及所有相关资源,以支持未来历史语言模型研究。
原文摘要 · Abstract (English)
We introduce TypewriterLM, a 7.24B History language model (LM) trained exclusively on English text predating 1913. Developing History LMs requires addressing challenges in data quality and availability, preventing temporal leakage, designing temporally consistent post-training pipelines, and constructing reliable evaluations. To address these issues, we construct TypewriterCorpus, a 54B-token historical corpus collected from diverse archival and linguistically annotated sources with extensive data cleaning and leakage mitigation procedures. Furthermore, we introduce lexically grounded instructing tuning, a post-training framework that constraints responses to remain directly grounded in historical source documents. Using this framework we construct two historical instruction tuning datasets: History-LIMA and History-SelfInstruct. To evaluate capability and temporal consistency, we introduce History-Event, a benchmark suite for evaluating competence, temporal grounding and data leakage. We release TypewriterLM and all associated resources to support future research on historical language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。