通过多智能体框架追溯大模型训练数据的演化路径,揭示数据间的系统关联。
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
- 构建多智能体系统自动重建数据演化图谱,追踪数据来源与演变关系。
- 发现数学类数据垂直精炼、通用数据横向聚合等结构性模式。
- 利用溯源信息生成更多样、去冗余的训练数据,提升数据质量。
后训练数据对大语言模型能力塑造至关重要,但现有数据常被视作孤立实体,忽视其背后的系统性关联。本文首次将‘数据溯源’引入大模型生态,提出一种自动化多智能体框架,用于重构数据集的演化图谱。大规模分析揭示:数学类数据呈现垂直精炼特征,通用领域数据则表现为水平聚合。同时发现广泛存在的结构性冗余问题,源于隐式数据交集,以及基准污染沿溯源路径传播的现象。为验证溯源分析的实际价值,我们基于重建的图谱设计了‘溯源感知的多样性导向数据集’,通过锚定上游根源进行指令采样,有效缓解下游同质化与隐藏冗余,生成更具多样性的后训练语料。此外,溯源分析可作为大规模数据生态中样本级比较的高效稳健拓扑替代方案。本研究推动后训练数据管理向更系统化、可控化的范式演进。
原文摘要 · Abstract (English)
Post-training data plays a pivotal role in shaping the capabilities of Large Language Models (LLMs), yet datasets are often treated as isolated artifacts, overlooking the systemic connections that underlie their evolution. To disentangle these complex relationships, we introduce the concept of \textbf{data lineage} to the LLM ecosystem and propose an automated multi-agent framework to reconstruct the evolutionary graph of dataset development. Through large-scale lineage analysis, we characterize domain-specific structural patterns, such as vertical refinement in math-oriented datasets and horizontal aggregation in general-domain corpora. Moreover, we uncover pervasive systemic issues, including \textit{structural redundancy} induced by implicit dataset intersections and the \textit{propagation of benchmark contamination} along lineage paths. To demonstrate the practical value of lineage analysis for data construction, we leverage the reconstructed lineage graph to create a \textit{lineage-aware diversity-oriented dataset}. By anchoring instruction sampling at upstream root sources, this approach mitigates downstream homogenization and hidden redundancy, yielding a more diverse post-training corpus. We further highlight lineage-centric analysis as an efficient and robust topological alternative to sample-level dataset comparison for large-scale data ecosystems. By grounding data construction in explicit lineage structures, our work advances post-training data curation toward a more systematic and controllable paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。