用年度模型追踪科学文本演变,揭示知识发展规律。
Towards understanding evolution of science through language model series
- 以整词为单位,逐年训练模型捕捉科学语言变迁。
- 在文献引用网络预测中表现优于现有方法。
- 适合关注科学演进与知识演化研究的学者。
我们提出AnnualBERT,一系列专为捕捉科学文本时间演化的语言模型。不同于主流的子词分词和单一通用模型范式,AnnualBERT采用整词作为标记,并基于170万篇截至2008年的arXiv论文全文从头预训练一个基础RoBERTa模型,再通过逐年增量训练构建系列模型。实验表明,AnnualBERT不仅在标准NLP任务上性能相当,还在领域特定任务及arXiv引文网络链接预测任务中达到当前最优表现。进一步通过探针任务量化模型随时间的表征学习与遗忘行为。该方法不仅能提升科学文本处理效果,还能揭示科学话语随时间的发展脉络。模型系列已开源至https://huggingface.co/jd445/AnnualBERTs。
原文摘要 · Abstract (English)
We introduce AnnualBERT, a series of language models designed specifically to capture the temporal evolution of scientific text. Deviating from the prevailing paradigms of subword tokenizations and "one model to rule them all", AnnualBERT adopts whole words as tokens and is composed of a base RoBERTa model pretrained from scratch on the full-text of 1.7 million arXiv papers published until 2008 and a collection of progressively trained models on arXiv papers at an annual basis. We demonstrate the effectiveness of AnnualBERT models by showing that they not only have comparable performances in standard tasks but also achieve state-of-the-art performances on domain-specific NLP tasks as well as link prediction tasks in the arXiv citation network. We then utilize probing tasks to quantify the models' behavior in terms of representation learning and forgetting as time progresses. Our approach enables the pretrained models to not only improve performances on scientific text processing tasks but also to provide insights into the development of scientific discourse over time. The series of the models is available at https://huggingface.co/jd445/AnnualBERTs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。