构建250年英文书籍时间结构语料库,助力大模型理解语言演变。
CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
- 基于古腾堡计划书籍构建带时间标注的英语语料库
- 发现语言模型难以捕捉意义随时间的变化,需改进训练与评估
- 适合研究语言演化、历史语义与跨时期情感分析的学者
大型语言模型(LLM)依赖网络爬取的社交媒体和多样化数据实现规模化运行。然而,现有语料库普遍缺乏长期时间结构,限制了模型对语言语义与规范演变的上下文理解及历时性差异的捕捉能力。为此,我们推出CHRONOBERG——一个涵盖250年英文书籍文本的时序结构语料库,源自古腾堡计划并附有多类时间标注。书籍的编辑性质使我们能通过时敏的情感三维度(VAD)分析量化词汇语义变化,并构建历史校准的情感词典,支持时间锚定的语义解读。基于该词典,我们揭示现代语言模型在识别歧视性语言和跨时期情绪判断中存在偏差。进一步实验表明,顺序训练于CHRONOBERG的语言模型难以编码意义的历时转移,凸显亟需引入时间感知的训练与评估流程。这使得CHRONOBERG成为研究语言变迁与时间泛化的重要可扩展资源。公开获取:https://huggingface.co/datasets/spaul25/Chronoberg;代码地址:https://github.com/paulsubarna/Chronoberg。
原文摘要 · Abstract (English)
Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with a variety of temporal annotations. First, the edited nature of books enables us to quantify lexical semantic change through time-sensitive Valence-Arousal-Dominance (VAD) analysis and to construct historically calibrated affective lexicons to support temporally grounded interpretation. With the lexicons at hand, we demonstrate a need for modern LLM-based tools to better situate their detection of discriminatory language and contextualization of sentiment across various time-periods. In fact, we show how language models trained sequentially on CHRONOBERG struggle to encode diachronic shifts in meaning, emphasizing the need for temporally aware training and evaluation pipelines, and positioning CHRONOBERG as a scalable resource for the study of linguistic change and temporal generalization. Disclaimer: This paper includes language and display of samples that could be offensive to readers. Open Access: Chronoberg is available publicly on HuggingFace at ( https://huggingface.co/datasets/spaul25/Chronoberg). Code is available at (https://github.com/paulsubarna/Chronoberg).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。