通过生成文本反推大模型训练数据构成,揭示其'数字基因'
LLMSurgeon: Diagnosing Data Mixture of Large Language Models

- 将数据混合恢复建模为带标签偏移假设的逆问题
- 在公开模型上恢复结果与真实数据混合高度一致
- 适合模型审计、数据溯源等可信AI研究者使用
大型语言模型(LLM)的预训练数据混合构成其“数字基因”,决定模型行为、能力与失效模式,但该组成通常不公开,导致事后审计困难。本文提出数据混合手术(DMS),仅基于目标模型生成文本,在预定义分类体系下估计其预训练语料的领域分布。我们提出LLMSurgeon框架,将问题视为标签偏移下的逆问题,不直接聚合分类器输出,而是估计校准的软混淆矩阵,并求解约束逆问题以纠正系统性领域混淆,恢复潜在混合先验。为评估,我们构建了LLMScan——一个基于开源模型且预训练混合透明的可验证评估套件。在LLMScan上,LLMSurgeon在固定协议下实现了高保真度的领域混合恢复。本工作提供了一种无需访问训练数据即可事后审计基础模型数字基因的实用方法。
原文摘要 · Abstract (English)
The pretraining data mixture of Large Language Models (LLMs) constitutes their "digital DNA", shaping model behaviors, capabilities, and failure modes. Yet this composition is rarely disclosed, making post-hoc auditing of data combination or provenance difficult. In this work, we formalize $\textbf{Data Mixture Surgery (DMS)}$: given only generated text from a target LLM, estimate the domain-level distribution of its pretraining corpus under a predefined taxonomy. We propose $\textbf{LLMSurgeon}$, a strong framework that casts DMS as an inverse problem under the label-shift assumption. Rather than directly aggregating classifier outputs, LLMSurgeon estimates a calibrated $\textit{soft}$ confusion matrix and solves a constrained inverse problem to correct systematic domain confusion and recover the latent mixture prior. To evaluate, we introduce $\textbf{LLMScan}$, a recipe-verifiable evaluation suite built from open-source LLMs with transparent pretraining mixtures. Across LLMScan, LLMSurgeon recovers domain mixtures with high fidelity under fixed protocols. Our work presents a practical, post-hoc approach for auditing the digital DNA of foundation models without access to their training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。