arXiv:2602.09003cs.AIcs.CL2026-02被引 2

提出分层数据管理框架,让大模型主动参与数据筛选与优化,提升训练效率和性能。

Data Science and Technology Towards AGI Part I: Tiered Data Management

  • 构建L0-L4五层数据体系,从原始数据到可验证知识,分级管理数据质量与用途。
  • 用大模型自动评分和编辑数据,实现数据在不同训练阶段的精准分配,提升训练效率。
  • 适合研究高效数据管理、大模型训练优化及可持续人工智能的开发者与研究人员。

人工智能的发展可视为数据驱动学习范式的演进,数据组织与利用方式的迭代持续推动模型能力提升。当前大语言模型研究主要依赖数据规模的单向扩展,正面临数据获取瓶颈、成本高企与训练效率下降等问题。本文认为,通用人工智能(AGI)发展已进入数据与模型协同演化的新阶段:模型主动引导数据管理,高质量数据反过来增强模型能力。为此,我们提出一种分层数据管理框架,支持异构学习目标与成本约束下的全周期大模型训练。该框架包含L0至L4共五个层级,覆盖从原始未标注资源到结构化可验证知识的数据形态。关键在于,大模型全程参与数据质量评估与内容编辑,实现跨层级数据精炼。各层级具有不同的数据特征、管理策略与训练角色,支持数据在预训练、中段训练与对齐阶段的战略性分配。框架在数据质量、获取成本与边际训练收益之间取得平衡,提供可扩展、可持续的数据管理方案。通过实证研究,我们从原始语料构建分层数据集,并在多个训练阶段应用。结果表明,基于层级感知的数据使用显著提升训练效率与模型表现。为促进后续研究,我们开源了分层数据集与处理工具。

原文摘要 · Abstract (English)

The development of artificial intelligence can be viewed as an evolution of data-driven learning paradigms, with successive shifts in data organization and utilization continuously driving advances in model capability. Current LLM research is dominated by a paradigm that relies heavily on unidirectional scaling of data size, increasingly encountering bottlenecks in data availability, acquisition cost, and training efficiency. In this work, we argue that the development of AGI is entering a new phase of data-model co-evolution, in which models actively guide data management while high-quality data, in turn, amplifies model capabilities. To implement this vision, we propose a tiered data management framework, designed to support the full LLM training lifecycle across heterogeneous learning objectives and cost constraints. Specifically, we introduce an L0-L4 tiered data management framework, ranging from raw uncurated resources to organized and verifiable knowledge. Importantly, LLMs are fully used in data management processes, such as quality scoring and content editing, to refine data across tiers. Each tier is characterized by distinct data properties, management strategies, and training roles, enabling data to be strategically allocated across LLM training stages, including pre-training, mid-training, and alignment. The framework balances data quality, acquisition cost, and marginal training benefit, providing a systematic approach to scalable and sustainable data management. We validate the effectiveness of the proposed framework through empirical studies, in which tiered datasets are constructed from raw corpora and used across multiple training phases. Experimental results demonstrate that tier-aware data utilization significantly improves training efficiency and model performance. To facilitate further research, we release our tiered datasets and processing tools to the community.

大模型数据管理训练效率分层架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。