arXiv:2511.23056cs.CL2025-11

用可解释的机器学习模型,精准判断英文历史文本的年代,还能看清语言演变规律。

Decoding the Past: Explainable Machine Learning Models for Dating Historical Texts

  • 融合五类特征,构建可解释的树模型预测文本年代。
  • 百年级预测准确率76.7%,十年级达26.1%,远超随机基准。
  • 模型可解释性强,揭示19世纪是语言演变的关键转折点。

准确判定历史文本年代对文化遗产整理与解读至关重要。本文采用可解释的、基于特征工程的树模型进行时间分类,整合压缩、词汇结构、可读性、新词检测和距离特征五类信息,预测跨越五个世纪的英文文本年代。对比分析表明,这些特征提供互补的时间信号,综合模型性能优于任一单一特征集。在大规模语料上,百年级预测准确率达76.7%,十年级为26.1%,显著高于随机基线(20% 和 2.3%)。在放宽精度要求下,百年级前2名准确率达96.0%,十年级前10名达85.8%。模型具备强排序能力,AUCROC最高达94.8%,AUPRC最高达83.3%,平均绝对误差分别为27年和30年。针对认证任务,在关键阈值(如1850–1900)附近,二分类模型准确率可达85%–98%。特征重要性分析显示,距离特征和词汇结构最有效,压缩特征提供补充信号。SHAP解释揭示系统性语言演化模式,19世纪成为多特征域的转折点。在Project Gutenberg数据集上的跨域评估显示准确率下降26.4个百分点,但树模型仍以高效与可解释性,成为神经网络的可扩展替代方案。

原文摘要 · Abstract (English)

Accurately dating historical texts is essential for organizing and interpreting cultural heritage collections. This article addresses temporal text classification using interpretable, feature-engineered tree-based machine learning models. We integrate five feature categories - compression-based, lexical structure, readability, neologism detection, and distance features - to predict the temporal origin of English texts spanning five centuries. Comparative analysis shows that these feature domains provide complementary temporal signals, with combined models outperforming any individual feature set. On a large-scale corpus, we achieve 76.7% accuracy for century-scale prediction and 26.1% for decade-scale classification, substantially above random baselines (20% and 2.3%). Under relaxed temporal precision, performance increases to 96.0% top-2 accuracy for centuries and 85.8% top-10 accuracy for decades. The final model exhibits strong ranking capabilities with AUCROC up to 94.8% and AUPRC up to 83.3%, and maintains controlled errors with mean absolute deviations of 27 years and 30 years, respectively. For authentication-style tasks, binary models around key thresholds (e.g., 1850-1900) reach 85-98% accuracy. Feature importance analysis identifies distance features and lexical structure as most informative, with compression-based features providing complementary signals. SHAP explainability reveals systematic linguistic evolution patterns, with the 19th century emerging as a pivot point across feature domains. Cross-dataset evaluation on Project Gutenberg highlights domain adaptation challenges, with accuracy dropping by 26.4 percentage points, yet the computational efficiency and interpretability of tree-based models still offer a scalable, explainable alternative to neural architectures.

文本年代可解释模型语言演化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。