用知识图谱结构组织海量网页数据,提升模型训练质量。
CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

- 构建三层异构知识图谱,实现内容、语义与跨域关联的系统化组织。
- 生成241.4亿token高质量语料库,支持跨领域检索与推理评测。
- 适合追求数据质量与结构化的AI研发团队使用。
大语言模型的持续演进对数据规模与质量提出更高要求,不同训练阶段需要定制化数据,系统化高质量语料组织变得不可或缺。现有语料构建流程仅生成扁平、无差别的文档集合,普遍缺乏系统的知识组织。我们提出Cortex,据我们所知首个将网络规模语料构建从扁平文档过滤升级为结构化知识组织的框架,通过本体语料图(OCG)实现三层异构结构:经质量优化的内容层、由大模型驱动自动演化的分层轻量本体层,以及支持任意分类层级跨域关联的对齐层。全面实验验证了Cortex的有效性。特别地,我们利用OCG构建CortexBench,一个跨领域搜索与推理基准,在八种前沿大模型上评估证实了质量优化、领域组织和跨域数据合成的有效性。我们将公开完整代码库、包含24.14亿token的精炼语料及其OCG,以及CortexBench。
原文摘要 · Abstract (English)
The continuous evolution of large language models drives escalating demands on data scale and quality, and as different training stages impose increasingly tailored data requirements, systematic organization of high-quality corpora becomes indispensable. Existing corpus construction pipelines confine the resulting corpora to flat, undifferentiated document collections, universally lacking systematic knowledge organization. We present Cortex, to our knowledge the first framework that elevates web-scale corpus construction from flat document filtering to structured knowledge organization through an Ontological Corpus Graph (OCG), a three-layer heterogeneous structure unifying a quality-refined content layer, a hierarchical lightweight ontology layer via LLM-driven automated evolution, and a cross-domain alignment layer enabling inter-domain association at arbitrary taxonomic resolution. Comprehensive experiments confirm the effectiveness of Cortex. In particular, we leverage the OCG to synthesize CortexBench, a cross-domain search-and-reasoning benchmark whose evaluation across eight frontier LLMs validates the effectiveness of quality refinement, domain organization, and cross-domain data synthesis. We will publicly release the complete codebase, a 24.14B-token refined corpus with its OCG, and CortexBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。