历史意大利语对大模型的挑战可拆解为编码与理解两部分,且可通过简单提示缓解。
How Surprising Is Historical Italian to Language Models? Tokenization Tax, Comprehension Tax, and a Simple Mitigation

- 将历史语言难度拆分为编码成本、预测不确定性等四维度进行诊断。
- 17世纪意大利语比现代语平均令人意外2.4倍,学术文本达3.2倍,但语义表示仍稳定。
- 加个时间上下文提示可降低60%的意外度,适合数字图书馆应用。
大语言模型在数字图书馆工作流中日益关键,但其处理历史语言的能力仍不清晰。本文提出诊断框架,将历史语言难度分解为四个维度:分词成本、预测不确定性(突兀性)、语义鲁棒性与上下文敏感性。我们在三个跨越三个世纪的数据集上评估:(1) 新整理的17世纪意大利语文本(1610–1689),源自原始图像;(2) 19世纪经典意大利语《失传的婚约》作为高暴露对照;(3) 18世纪俄国民间印刷书作为对比正字法压力测试。结果揭示编码成本与理解能力存在显著分离:俄语和早期现代意大利语分词开销相当(通胀25–30%),但预测难度差异显著。17世纪意大利语平均比现代语突兀2.4倍,学术文本达3.2倍,而俄语仅轻微上升。但预测不确定性不意味表征退化:嵌入相似性在所有数据集均保持>0.85,说明模型仍能表征历史语义。最后,我们证明加入最小时间上下文提示可降低约60%的历史突兀性,提供一种简单、模型无关的缓解策略。表明尽管历史文本有持续编码成本,数字图书馆可在语义检索任务中安全使用大模型,只需对生成类应用做适配。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly critical to digital library workflows, yet their ability to process historical language remains poorly understood. Historical difficulty is typically treated as a monolithic barrier, conflating orthographic variation, linguistic distance, and pretraining exposure. In this paper, we propose a diagnostic framework that decomposes this difficulty into four distinct dimensions: tokenization cost, predictive uncertainty (surprisal), semantic robustness, and context sensitivity. We evaluate this framework on three datasets spanning three centuries: (1) a newly curated corpus of 17th-century Italian texts (1610-1689) digitized from original page images; (2) canonical 19th-century Italian "I Promessi Sposi" serving as a high-exposure control; and (3) 18th-century Russian civil print books as a contrastive orthographic stress test. Our results reveal a distinct dissociation between encoding cost and comprehension. While Russian and early modern Italian incur comparable tokenization penalties (25-30% inflation), their predictive difficulty diverges sharply. 17th-century Italian is on average 2.4 times more surprising than its modern equivalent - with academic prose reaching 3.2 times - whereas Russian shows only a modest increase. But predictive uncertainty does not imply representational degradation: embedding similarity remains robust (> 0.85) across all datasets, confirming that models can represent historical meaning even when generation is unstable. Finally, we demonstrate that a minimal temporal context prompt reduces historical surprisal by approximately 60%, offering a simple, model-agnostic mitigation. These findings suggest that while historical text imposes a consistent encoding tax, digital libraries can safely deploy LLMs for semantic retrieval tasks, provided that generative applications are carefully adapted.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。