arXiv:2505.16260cs.LG2025-05被引 1

训练数据分布对大小模型的影响高度一致,可借小模型推断大模型表现。

Small-to-Large Generalization: Data Influences Models Consistently Across Scale

  • 通过小模型推断大模型在不同数据下的行为,验证数据影响的一致性。
  • 发现小模型与大模型在不同数据分布下的预测相关性高,达0.8以上。
  • 适用于数据筛选和归因分析,尤其适合资源有限的研究者。

训练数据分布对模型行为有显著影响。然而,在大规模模型中,由于训练成本高,精确分析数据变化对预测的影响十分困难。当前做法是依赖低成本的小规模代理模型进行外推。但小模型与大模型对数据变化的响应并不相同。本研究发现,小规模与大规模语言模型在不同训练数据下的预测结果具有高度相关性(相关系数普遍超过0.8)。基于此,我们进一步分析了代理模型规模对两类下游任务——数据归因与数据集选择——的有效性,证明使用小模型进行推断在跨规模场景中具备可靠参考价值。

原文摘要 · Abstract (English)

Choice of training data distribution greatly influences model behavior. Yet, in large-scale settings, precisely characterizing how changes in training data affects predictions is often difficult due to model training costs. Current practice is to instead extrapolate from scaled down, inexpensive-to-train proxy models. However, changes in data do not influence smaller and larger models identically. Therefore, understanding how choice of data affects large-scale models raises the question: how does training data distribution influence model behavior across compute scale? We find that small- and large-scale language model predictions (generally) do highly correlate across choice of training data. Equipped with these findings, we characterize how proxy scale affects effectiveness in two downstream proxy model applications: data attribution and dataset selection.

模型泛化数据影响代理模型语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。