arXiv:2502.03656cs.CVcs.AI2025-02被引 2

用不到9%的数据量,训练出与全量数据效果相近的超分辨率模型。

A Study in Dataset Distillation for Image Super-Resolution

  • 从像素和潜在空间两个角度设计压缩方法,保留原始数据的训练特性。
  • 仅需原数据8.88%的样本量,就能达到接近全量训练的重建精度。
  • 为节省内存和计算资源的生成式图像修复提供实用方案。

数据蒸馏旨在将大规模数据集压缩为紧凑且信息丰富的子集,同时保持原始数据的训练行为。尽管该概念在分类任务中已受关注,但在图像超分辨率(SR)领域仍基本未被探索。本文首次系统研究了针对SR的数据蒸馏,评估了像素空间与潜在空间两种范式。实验表明,仅占原数据8.88%的蒸馏数据集,即可训练出在重建保真度上几乎等同于全量数据训练的SR模型。此外,我们分析了初始化策略与蒸馏目标对效率、收敛速度及视觉质量的影响。研究结果验证了SR数据蒸馏的可行性,并为构建内存与计算高效生成式恢复模型提供了基础洞见。

原文摘要 · Abstract (English)

Dataset distillation aims to compress large datasets into compact yet highly informative subsets that preserve the training behavior of the original data. While this concept has gained traction in classification, its potential for image Super-Resolution (SR) remains largely untapped. In this work, we conduct the first systematic study of dataset distillation for SR, evaluating both pixel- and latent-space formulations. We show that a distilled dataset, occupying only 8.88% of the original size, can train SR models that retain nearly the same reconstruction fidelity as those trained on full datasets. Furthermore, we analyze how initialization strategies and distillation objectives affect efficiency, convergence, and visual quality. Our findings highlight the feasibility of SR dataset distillation and establish foundational insights for memory- and compute-efficient generative restoration models.

图像超分数据蒸馏模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。