用合成数据探针解析数据如何影响大模型性能
Position: Let's Develop Data Probes to Fundamentally Understand How Data Affects LLM Performance

- 设计随机生成的合成数据序列作为探针
- 通过探针观察模型在不同阶段的表现差异
- 适合研究数据本质与模型行为关系的学者
数据是大语言模型(LLM)的核心。然而,当前对特定数据为何在训练、微调、对齐、上下文学习等阶段有效仍缺乏系统理解。现有方法依赖大规模公共数据集的反复实验,获取经验性数据筛选规则,计算成本高且缺乏理论依据。本文倡导发展系统性方法,从定义良好的随机过程生成合成序列,用于揭示数据特征如何驱动模型行为。这类序列称为数据探针。通过观测模型在探针上的表现,可系统研究数据特性对模型性能、泛化能力与鲁棒性的影响。探针具备可分析的统计性质,如典型集等理论概念可用于解释模型行为。该方法为超越经验启发、深入理解数据在模型训练与推理中的基础作用提供新路径。
原文摘要 · Abstract (English)
Data is fundamental to large language models (LLMs). However, understanding of what makes certain data useful for different stages of an LLM workflow, including training, tuning, alignment, in-context learning, etc., and why, remains an open question. Current approaches rely heavily on extensive experimentation with large public datasets to obtain empirical heuristics for data filtering and dataset construction. These approaches are compute intensive and lack a principled way of understanding the essence of how specific data characteristics drive LLM behavior. In this position paper, we advocate for the need of developing systematic methodologies for generating synthetic sequences from appropriately defined random processes, with the goal that these sequences can reveal useful characteristics when they are used in one or multiple stages of the LLM workflow. We refer to such sequences as data probes. By observing LLM behavior on data probes, researchers can systematically conduct studies on how data characteristics influence model performance, generalization, and robustness. The probing sequences exhibit statistical properties that can be viewed using theoretical concepts, such as typical sets, which are generalized to describe the behaviors of LLMs. This data-probe approach provides a pathway for uncovering foundational insights into the role of data in LLM training and inference, beyond empirical heuristics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。