arXiv:2511.19274cs.CV2025-11AAAI

用扩散模型重建误差评估数据重要性,选50%数据效果接近全量训练。

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

  • 通过部分逆向去噪的重建偏差估算数据似然,理论基于扩散过程的变分下界。
  • 在ImageNet上仅用50%数据即逼近全量训练性能,优于现有基线方法。
  • 适合关注数据筛选机制、分布建模与高效训练的研究者使用。

现有核心集选择方法多依赖训练动态或模型不确定性的启发式评分信号,缺乏对数据似然的显式建模,可能难以捕捉支撑有效训练的细微分布结构。本文提出一种新方法,利用扩散模型通过部分逆向去噪引发的重建偏差来估计数据似然。我们建立了重建误差与数据似然之间的形式化联系,基于马尔可夫扩散过程的证据下界(ELBO),从而实现一种原理严谨、分布感知的数据选择评分准则。同时引入高效的信道理论方法确定最优重建时间步,确保偏差信号可靠反映底层数据似然。ImageNet上的大量实验表明,重建偏差作为评分标准表现优异,在不同选取比例下持续超越现有基线,并在仅使用50%数据时接近全量数据训练效果。进一步分析显示,该评分具有似然感知特性,揭示了数据分布特征与模型学习偏好之间的内在关联。

原文摘要 · Abstract (English)

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed subset from capturing subtle yet critical distributional structures that underpin effective model training. In this work, we propose a novel, theoretically grounded approach that leverages diffusion models to estimate data likelihood via reconstruction deviation induced by partial reverse denoising. Specifically, we establish a formal connection between reconstruction error and data likelihood, grounded in the Evidence Lower Bound (ELBO) of Markovian diffusion processes, thereby enabling a principled, distribution-aware scoring criterion for data selection. Complementarily, we introduce an efficient information-theoretic method to identify the optimal reconstruction timestep, ensuring that the deviation provides a reliable signal indicative of underlying data likelihood. Extensive experiments on ImageNet demonstrate that reconstruction deviation offers an effective scoring criterion, consistently outperforming existing baselines across selection ratios, and closely matching full-data training using only 50% of the data. Further analysis shows that the likelihood-informed nature of our score reveals informative insights in data selection, shedding light on the interplay between data distributional characteristics and model learning preferences.

核心集选择扩散模型数据效率分布建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。