arXiv:2502.01980cs.LGcs.AI2025-02ICML被引 3

用扩散模型主动生成罕见数据,提升模型泛化能力

Generative Data Mining with Longtail-Guided Diffusion

  • 用可微分的不确定性信号识别罕见输入,无需修改原模型
  • 通过长尾引导生成新训练数据,在多个图像分类任务上显著提优
  • 生成数据具语义意义,可被视觉语言模型分析并解释模型缺陷

预测模型在部署后可能面临诸多未知挑战。现有做法多为被动循环:部署、数据挖掘、再训练。本文提出主动的长尾发现机制,在训练阶段即模拟生成额外数据。我们设计通用的基于模型的长尾信号,包括一种可微分的单次前向传播不确定性度量,不改变模型参数或预测性能,但能标记稀有或困难样本。利用这些信号作为指导,通过潜在扩散模型生成新训练数据,该过程称为长尾引导(LTG)。关键在于,LTG无需重训扩散模型或预测模型,也不必暴露预测模型于中间扩散状态。由LTG生成的数据具有语义合理性,在多个图像分类基准上带来显著泛化提升,并可通过视觉语言模型进行分析,主动发现、文本解释并修复部署模型中的概念性缺陷。

原文摘要 · Abstract (English)

It is difficult to anticipate the myriad challenges that a predictive model will encounter once deployed. Common practice entails a reactive, cyclical approach: model deployment, data mining, and retraining. We instead develop a proactive longtail discovery process by imagining additional data during training. In particular, we develop general model-based longtail signals, including a differentiable, single forward pass formulation of epistemic uncertainty that does not impact model parameters or predictive performance but can flag rare or hard inputs. We leverage these signals as guidance to generate additional training data from a latent diffusion model in a process we call Longtail Guidance (LTG). Crucially, we can perform LTG without retraining the diffusion model or the predictive model, and we do not need to expose the predictive model to intermediate diffusion states. Data generated by LTG exhibit semantically meaningful variation, yield significant generalization improvements on numerous image classification benchmarks, and can be analyzed by a VLM to proactively discover, textually explain, and address conceptual gaps in a deployed predictive model.

扩散模型长尾数据生成式挖掘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。