构建41种语言的跨时语料库,助力语义演变研究
DHPLT: large-scale multilingual diachronic corpora and word representations for semantic change modelling
- 基于网页爬取数据,按时间戳划分2011-2015、2020-2021、2024至今三时段语料
- 每语言每时段100万文档,提供词嵌入与词汇替换数据
- 支持多语言语义变化建模,适合自然语言处理与历史语言学研究者
本文介绍DHPLT,一个涵盖41种不同语言的大规模跨时语料库资源。该语料库基于网络爬取的HPLT数据集,利用网页抓取时间戳作为文档创建时间的近似信号。数据覆盖三个时期:2011–2015年、2020–2021年及2024年至今,每语言每个时期均包含约100万份文档。我们还提供了预计算的词类型与词元嵌入,以及选定目标词的词汇替换信息;同时开放数据,允许研究者使用相同数据集自定义目标词。DHPLT旨在填补当前语义变化建模中多语言跨时语料库的不足(现有资源仅覆盖十余种高资源语言),为该领域开辟多种新实验路径。所有资源均可在https://data.hplt-project.org/three/diachronic/获取,按语言分类。
原文摘要 · Abstract (English)
In this resource paper, we present DHPLT, an open collection of diachronic corpora in 41 diverse languages. DHPLT is based on the web-crawled HPLT datasets; we use web crawl timestamps as the approximate signal of document creation time. The collection covers three time periods: 2011-2015, 2020-2021 and 2024-present (1 million documents per time period for each language). We additionally provide pre-computed word type and token embeddings and lexical substitutions for our chosen target words, while at the same time leaving it open for the other researchers to come up with their own target words using the same datasets. DHPLT aims at filling in the current lack of multilingual diachronic corpora for semantic change modelling (beyond a dozen of high-resource languages). It opens the way for a variety of new experimental setups in this field. All the resources described in this paper are available at https://data.hplt-project.org/three/diachronic/, sorted by language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。