arXiv:2510.21762cs.CLcs.DL2025-10

构建83万条多语言科学论文段落分类数据集,支持文献挖掘模型训练。

A Multi-lingual Dataset of Classified Paragraphs from Open Access Scientific Publications

  • 从开放获取论文中提取段落,按四类主题分类
  • 覆盖英法及多种欧洲语言,含语言与领域标签
  • 适合做科学文献命名实体识别与多语言分类任务

我们构建了一个包含83.3万条段落的数据集,这些段落来自采用CC-BY许可的科学出版物,按四类主题分类:致谢、数据提及、软件/代码提及和临床试验提及。段落主要为英语和法语,还包含其他欧洲语言。每条段落均标注了语言(使用fastText)和科学领域(来自OpenAlex)。该数据集源自法国开放科学监测器语料库,经GROBID处理生成,可用于训练文本分类模型及开发科学文献命名实体识别系统。数据集已公开发布于HuggingFace,网址为https://doi.org/10.57967/hf/6679,采用CC-BY许可。

原文摘要 · Abstract (English)

We present a dataset of 833k paragraphs extracted from CC-BY licensed scientific publications, classified into four categories: acknowledgments, data mentions, software/code mentions, and clinical trial mentions. The paragraphs are primarily in English and French, with additional European languages represented. Each paragraph is annotated with language identification (using fastText) and scientific domain (from OpenAlex). This dataset, derived from the French Open Science Monitor corpus and processed using GROBID, enables training of text classification models and development of named entity recognition systems for scientific literature mining. The dataset is publicly available on HuggingFace https://doi.org/10.57967/hf/6679 under a CC-BY license.

科学文献多语言数据集文本分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。