用扩散模型生成视频并高效筛选关键片段,大幅降低训练成本。
Video Dataset Condensation with Diffusion Models
- 用视频扩散模型一次性生成合成视频,减少重复计算。
- 提出VST-UNet和无训练聚类TAC-DT,提升代表性视频选择效率。
- 在4个数据集上性能超现有方法最高10.61%,适合视频压缩与训练优化场景。
近年来,数据集规模迅速扩大和深度学习模型复杂度上升,导致计算资源需求激增,涵盖存储与训练两方面。数据集浓缩成为应对挑战的有前景方案,通过生成紧凑的合成数据集来保留原始大数据集的关键信息。然而,现有方法在视频领域表现受限。本文聚焦视频数据集浓缩,首先利用视频扩散模型生成合成视频,因仅需一次生成,显著降低计算开销。接着,提出视频时空U-Net(VST-UNet),用于选取多样且信息丰富的视频子集,有效捕捉原数据特征。为进一步提升效率,引入无需训练的时序感知聚类方法TAC-DT,实现代表性视频的高效筛选。在四个基准数据集上进行大量实验,验证了该方法的有效性,性能最高比当前最优提升10.61%。本方法在所有数据集上均优于现有方法,建立了视频数据集浓缩的新基准。
原文摘要 · Abstract (English)
In recent years, the rapid expansion of dataset sizes and the increasing complexity of deep learning models have significantly escalated the demand for computational resources, both for data storage and model training. Dataset distillation has emerged as a promising solution to address this challenge by generating a compact synthetic dataset that retains the essential information from a large real dataset. However, existing methods often suffer from limited performance, particularly in the video domain. In this paper, we focus on video dataset distillation. We begin by employing a video diffusion model to generate synthetic videos. Since the videos are generated only once, this significantly reduces computational costs. Next, we introduce the Video Spatio-Temporal U-Net (VST-UNet), a model designed to select a diverse and informative subset of videos that effectively captures the characteristics of the original dataset. To further optimize computational efficiency, we explore a training-free clustering algorithm, Temporal-Aware Cluster-based Distillation (TAC-DT), to select representative videos without requiring additional training overhead. We validate the effectiveness of our approach through extensive experiments on four benchmark datasets, demonstrating performance improvements of up to \(10.61\%\) over the state-of-the-art. Our method consistently outperforms existing approaches across all datasets, establishing a new benchmark for video dataset distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。