构建两个维基百科量值数据集,助力自动提取科学文本中的定量信息。
Wiki-Quantities and Wiki-Measurements: Datasets of Quantities and their Measurement Context from Wikipedia
- 基于维基百科和维基数据构建量值标注数据集
- 含120万条量值与3.8万条带测量上下文的量值
- 适用于科研文献自动信息抽取,代码开源可复现
为应对大量学术出版物,越来越多研究者采用自然语言处理技术自动提取感兴趣的数据。尤其在自然科学与工程领域,数据多为定量形式,但缺乏用于识别文本中量值及其测量上下文的数据集。为此,我们基于维基百科和维基数据构建了两个大规模数据集:Wiki-Quantities 包含超过120万条英文维基百科中的标注量值;Wiki-Measurements 包含38,738条标注量值,附带对应的被测实体、属性及可选限定词。对各数据集随机抽样100条进行人工验证,正确率分别为100%和84%-94%。这些数据集可用于测量信息抽取的流水线方法,先识别量值,再解析其测量上下文。为支持使用新版本维基百科或维基数据复现本工作,我们公开了全部数据生成代码。
原文摘要 · Abstract (English)
To cope with the large number of publications, more and more researchers are automatically extracting data of interest using natural language processing methods based on supervised learning. Much data, especially in the natural and engineering sciences, is quantitative, but there is a lack of datasets for identifying quantities and their context in text. To address this issue, we present two large datasets based on Wikipedia and Wikidata: Wiki-Quantities is a dataset consisting of over 1.2 million annotated quantities in the English-language Wikipedia. Wiki-Measurements is a dataset of 38,738 annotated quantities in the English-language Wikipedia along with their respective measured entity, property, and optional qualifiers. Manual validation of 100 samples each of Wiki-Quantities and Wiki-Measurements found 100% and 84-94% correct, respectively. The datasets can be used in pipeline approaches to measurement extraction, where quantities are first identified and then their measurement context. To allow reproduction of this work using newer or different versions of Wikipedia and Wikidata, we publish the code used to create the datasets along with the data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。