arXiv:2507.22360cs.CVcs.AI2025-07被引 2

用扩散模型压缩视频数据,用极少帧达到接近全量训练效果

GVD: Guiding Video Diffusion Model for Scalable Video Distillation

  • 首次提出基于扩散模型的视频蒸馏方法,联合优化时空特征
  • 在MiniUCF仅用1.98%帧达原性能78.29%,HMDB51用3.30%帧达73.83%
  • 支持高分辨率、高实例数训练,计算开销低,适合资源受限场景

为应对大规模视频数据带来的计算与存储压力,视频数据集蒸馏旨在以更小的数据集保留空间与时间信息,使在蒸馏数据上训练的模型性能接近使用全部数据时的表现。本文提出GVD:Guiding Video Diffusion,首个基于扩散模型的视频蒸馏方法。GVD联合蒸馏时空特征,在多样动作下实现高保真视频生成,并捕捉关键运动信息。在MiniUCF与HMDB51数据集上,5、10、20个实例每类(IPC)条件下,本方法显著优于现有最先进方法。具体而言,在MiniUCF上仅使用1.98%的总帧数即达到原数据集78.29%的性能;在HMDB51上仅用3.30%帧数即达到73.83%性能。实验表明,GVD不仅达到当前最优表现,还能生成更高分辨率视频,支持更高IPC,且计算成本几乎不变。

原文摘要 · Abstract (English)

To address the larger computation and storage requirements associated with large video datasets, video dataset distillation aims to capture spatial and temporal information in a significantly smaller dataset, such that training on the distilled data has comparable performance to training on all of the data. We propose GVD: Guiding Video Diffusion, the first diffusion-based video distillation method. GVD jointly distills spatial and temporal features, ensuring high-fidelity video generation across diverse actions while capturing essential motion information. Our method's diverse yet representative distillations significantly outperform previous state-of-the-art approaches on the MiniUCF and HMDB51 datasets across 5, 10, and 20 Instances Per Class (IPC). Specifically, our method achieves 78.29 percent of the original dataset's performance using only 1.98 percent of the total number of frames in MiniUCF. Additionally, it reaches 73.83 percent of the performance with just 3.30 percent of the frames in HMDB51. Experimental results across benchmark video datasets demonstrate that GVD not only achieves state-of-the-art performance but can also generate higher resolution videos and higher IPC without significantly increasing computational cost.

视频蒸馏扩散模型数据压缩高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。