用高效预训练发现语言随时间演变的规律,速度快且精准。
Pretraining Language Models for Diachronic Linguistic Change Discovery
- 在五段各1000万词的时间分段语料上,用高效预训练构建模型
- 模型比微调基线更快训练,且更准确反映语料的历史分期
- 适合历史语言学等领域快速发现词汇、语法等演变现象
大型语言模型在科学发现中展现出潜力,尤其在历史语言学和文学研究中应用日益广泛。这些领域常依赖体裁或时间分期进行论证。尽管已有方法通过微调或模型编辑限制领域,但真正可靠的保障仍是领域限定的预训练——这通常成本高昂。本文展示,高效预训练技术可在无法人工检查但又不足以支持传统大模型训练的大规模语料上取得良好效果。我们构建了一个包含五个1000万词时间片段的时序分段数据集,训练了两组五模型体系:高效预训练模型与参数量为80亿的Llama3-8B高效微调模型。结果表明,预训练模型训练速度更快,且更忠实于语料的历史划分。强调速度与精度而非跨时代全面性,为假设发现与验证提供了新路径。以历时语言学为测试场景,该方法成功检测到大规模词汇变化、非词汇性(语法与形态)变化以及词义的出现与消亡。本文提供一个可直接使用的流程,仅需少量调整即可推广至其他领域。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown potential as tools for scientific discovery. This has engendered growing interest in their use in humanistic disciplines, such as historical linguistics and literary studies. These fields often construct arguments on the basis of delineations like genre, or more inflexibly, time period. Although efforts have been made to restrict inference to specific domains via fine-tuning or model editing, we posit that the only true guarantee is domain-restricted pretraining -- typically, a data- and compute-expensive proposition. We show that efficient pretraining techniques can produce useful models over corpora too large for easy manual inspection but too small for "typical" LLM approaches. We employ a novel date-attribution pipeline in order to obtain a temporally-segmented dataset of five 10-million-word slices. We train two corresponding five-model batteries over these corpus segments, efficient pretraining and Llama3-8B parameter efficiently finetuned. We find that the pretrained models are faster to train than the finetuned baselines and that they better respect the historical divisions of our corpus. Emphasizing speed and precision over a-historical comprehensiveness enables a number of novel approaches to hypothesis discovery and testing in our target fields. Taking up diachronic linguistics as a testbed, we show that our method enables the detection of a diverse set of phenomena, including en masse lexical change, non-lexical (grammatical and morphological) change, and word sense introduction/obsolescence. We provide a ready-to-use pipeline that allows extension of our approach to other target fields with only minimal adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。