在模型不可见且资源有限时,仍可高效进行数据影响分析。
Exploring Training Data Attribution under Limited Access Constraints
- 用代理模型解决无完整模型访问的问题。
- 未预训练的模型也能提供有信息量的归因结果。
- 适合资源受限场景下的数据溯源与质量评估。
训练数据归因(TDA)对于理解单个训练数据对模型预测的影响至关重要。基于梯度的TDA方法,尤其是影响力函数,因其优异性能被广泛用于数据筛选、清洗、数据经济和事实追溯。然而,在商业模型不公开、计算资源有限的真实场景中,现有TDA方法受限于对完整模型访问的需求和高计算成本,难以推广。本文系统研究了在不同访问与资源约束下TDA方法的可行性,通过设计合适的代理模型等方案探索其适用性。我们发现,即使模型未在目标数据集上预训练,其获取的归因分数在多种任务中仍具有信息量,这对计算资源受限的场景尤为有用。研究结果为真实环境中TDA的部署提供了实用指导,旨在提升其可行性和效率。
原文摘要 · Abstract (English)
Training data attribution (TDA) plays a critical role in understanding the influence of individual training data points on model predictions. Gradient-based TDA methods, popularized by \textit{influence function} for their superior performance, have been widely applied in data selection, data cleaning, data economics, and fact tracing. However, in real-world scenarios where commercial models are not publicly accessible and computational resources are limited, existing TDA methods are often constrained by their reliance on full model access and high computational costs. This poses significant challenges to the broader adoption of TDA in practical applications. In this work, we present a systematic study of TDA methods under various access and resource constraints. We investigate the feasibility of performing TDA under varying levels of access constraints by leveraging appropriately designed solutions such as proxy models. Besides, we demonstrate that attribution scores obtained from models without prior training on the target dataset remain informative across a range of tasks, which is useful for scenarios where computational resources are limited. Our findings provide practical guidance for deploying TDA in real-world environments, aiming to improve feasibility and efficiency under limited access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。