arXiv:2601.22276cs.LGcs.CV2026-01

无需重新训练,快速评估文本生成图像模型中数据贡献者的重要性。

SurrogateSHAP: Training-Free Contributor Attribution for Text-to-Image (T2I) Models

  • 用预训练模型推理替代重训练,大幅降低计算开销。
  • 在三个数据集上均优于现有方法,准确识别关键贡献数据。
  • 适合数据市场、版权分配及医疗图像安全审计场景。

随着文本到图像(T2I)扩散模型在真实创作流程中的广泛应用,为数据提供者建立公平的贡献估值框架至关重要。尽管谢林值(Shapley value)在理论上具有合理性,但其面临双重计算瓶颈:(i) 每次采样数据贡献者子集都需要耗费巨大成本的模型重训练;(ii) 由于贡献者间交互,需估算大量组合子集以估计边际贡献。为此,我们提出 SurrogateSHAP,一种无需重训练的框架,通过从预训练模型中进行推理来近似昂贵的重训练过程。为进一步提升效率,我们采用梯度提升树逼近效用函数,并从树模型中解析推导出谢林值。我们在三项不同任务中评估了 SurrogateSHAP:(i) DDPM-CFG 在 CIFAR-20 上的图像质量;(ii) Stable Diffusion 在后印象派艺术作品上的美学评分;(iii) FLUX.1 在 Fashion-Product 数据上的产品多样性。在各类设置下,SurrogateSHAP 均显著优于先前方法,同时大幅减少计算开销,稳定识别出多个效用指标下的关键贡献者。最后,我们证明 SurrogateSHAP 能有效定位临床图像中导致虚假相关性的数据来源,为审计高风险生成模型提供可扩展路径。

原文摘要 · Abstract (English)

As Text-to-Image (T2I) diffusion models are increasingly used in real-world creative workflows, a principled framework for valuing contributors who provide a collection of data is essential for fair compensation and sustainable data marketplaces. While the Shapley value offers a theoretically grounded approach to attribution, it faces a dual computational bottleneck: (i) the prohibitive cost of exhaustive model retraining for each sampled subset of players (i.e., data contributors) and (ii) the combinatorial number of subsets needed to estimate marginal contributions due to contributor interactions. To this end, we propose SurrogateSHAP, a retraining-free framework that approximates the expensive retraining game through inference from a pretrained model. To further improve efficiency, we employ a gradient-boosted tree to approximate the utility function and derive Shapley values analytically from the tree-based model. We evaluate SurrogateSHAP across three diverse attribution tasks: (i) image quality for DDPM-CFG on CIFAR-20, (ii) aesthetics for Stable Diffusion on Post-Impressionist artworks, and (iii) product diversity for FLUX.1 on Fashion-Product data. Across settings, SurrogateSHAP outperforms prior methods while substantially reducing computational overhead, consistently identifying influential contributors across multiple utility metrics. Finally, we demonstrate that SurrogateSHAP effectively localizes data sources responsible for spurious correlations in clinical images, providing a scalable path toward auditing safety-critical generative models.

生成模型数据贡献公平性审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。