arXiv:2603.14712cs.CLcs.LG2026-03被引 2

构建智能数据系统,让大模型训练更高效、更自适应。

Towards Next-Generation LLM Training: From the Data-Centric Perspective

  • 用智能代理自动构建可复用的数据流水线,减少人工错误。
  • 训练中动态选择、混合和重加权数据,提升利用效率。
  • 适合关注数据工程与训练优化的研究者和工程师。

大型语言模型在众多任务和领域中表现出色,数据在其成功中起核心作用。然而,大规模模型训练所需数据的准备与有效利用仍是主要瓶颈。当前实践中,训练数据常通过临时脚本构建,缺乏成熟、基于智能体的数据准备系统,无法自动构建稳健且可复用的工作流,导致数据科学家陷入重复性、易出错的工程工作。此外,数据收集后通常被整体使用,缺乏系统性的数据筛选、混合优化或重加权机制。为此,我们提出两个互补方向:一是构建基于智能体的自动化数据准备系统,支持工作流自动生成与可扩展数据管理;二是建立统一的数据-模型交互训练系统,实现训练过程中数据的动态选择、混合与重加权,以提升训练效率、适应性与性能感知能力。最后,我们讨论了现存挑战,并展望未来研究与系统发展的前景。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable performance across a wide range of tasks and domains, with data playing a central role in enabling these advances. Despite this success, the preparation and effective utilization of the massive datasets required for LLM training remain major bottlenecks. In current practice, LLM training data is often constructed using ad hoc scripts, and there is still a lack of mature, agent-based data preparation systems that can automatically construct robust and reusable data workflows, thereby freeing data scientists from repetitive and error-prone engineering efforts. Moreover, once collected, datasets are often consumed largely in their entirety during training, without systematic mechanisms for data selection, mixture optimization, or reweighting. To address these limitations, we advocate two complementary research directions. First, we propose building a robust, agent-based automatic data preparation system that supports automated workflow construction and scalable data management. Second, we argue for a unified data-model interaction training system in which data is dynamically selected, mixed, and reweighted throughout the training process, enabling more efficient, adaptive, and performance-aware data utilization. Finally, we discuss the remaining challenges and outline promising directions for future research and system development.

数据工程LLM训练智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。