构建可复现的化学文献语料库,支持下游检索与文本挖掘。
Lit2Vec: A Reproducible Workflow for Building a Legally Screened Chemistry Corpus from S2ORC for Downstream Retrieval and Text Mining
- 基于元数据筛选法律合规的化学论文,自动构建结构化语料
- 包含58万篇全文文章、段落嵌入和18个领域标注
- 提供完整代码与流程,可复现语料构建过程
我们提出Lit2Vec,一个可复现的工作流,用于从学术开放研究语料库(S2ORC)中构建并验证化学领域语料库。通过保守的元数据许可筛查策略,共收集582,683篇化学领域的全文研究论文,具备结构化全文、分段切块、段落级嵌入(使用intfloat/e5-large-v2模型生成)以及记录级元数据(包括摘要和授权信息)。为支持下游检索与文本挖掘任务,该语料库还额外增加了机器生成的简要摘要和覆盖18个化学子领域的多标签分类标注。许可筛选利用Unpaywall、OpenAlex和Crossref的元数据完成,并对语料库进行了模式合规性、嵌入可复现性、文本质量和元数据完整性等方面的验证。主要贡献在于一套可复现的语料构建与验证工作流,及其配套的模式和可复现性资源。发布的材料包括代码、重建流程、模式定义、元数据/来源追踪文件及验证结果,可基于固定公开上游资源复现语料库。源文本及衍生表示的公共再分发不在本次发布范围内。研究人员可通过使用公开的上游数据集与元数据服务,复现该工作流。
原文摘要 · Abstract (English)
We present Lit2Vec, a reproducible workflow for constructing and validating a chemistry corpus from the Semantic Scholar Open Research Corpus using conservative, metadata-based license screening. Using this workflow, we assembled an internal study corpus of 582,683 chemistry-specific full-text research articles with structured full text, token-aware paragraph chunks, paragraph-level embeddings generated with the intfloat/e5-large-v2 model, and record-level metadata including abstracts and licensing information. To support downstream retrieval and text-mining use cases, an eligible subset of the corpus was additionally enriched with machine-generated brief summaries and multi-label subfield annotations spanning 18 chemistry domains. Licensing was screened using metadata from Unpaywall, OpenAlex, and Crossref, and the resulting corpus was technically validated for schema compliance, embedding reproducibility, text quality, and metadata completeness. The primary contribution of this work is a reproducible workflow for corpus construction and validation, together with its associated schema and reproducibility resources. The released materials include the code, reconstruction workflow, schema, metadata/provenance artifacts, and validation outputs needed to reproduce the corpus from pinned public upstream resources. Public redistribution of source-derived text and broad text-derived representations is outside the scope of the general release. Researchers can reproduce the workflow by using the released pipeline with publicly available upstream datasets and metadata services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。