arXiv:2505.20694cs.CVcs.LG2025-05被引 2

用时间显著性指导视频数据蒸馏,提升压缩效率与运动保真度。

Temporal Saliency-Guided Distillation: A Scalable Framework for Distilling Video Datasets

  • 基于预训练模型优化合成视频,引入时间显著性过滤冗余帧
  • 在多个视频基准上达到当前最佳性能,逼近真实数据效果
  • 适合大规模视频数据压缩与高效训练场景

数据蒸馏(DD)已成为数据集压缩的强大范式,能够生成紧凑的替代数据集以逼近大规模数据的训练效果。尽管图像数据蒸馏已取得显著进展,但将DD扩展至视频领域仍面临高维与时序复杂性的挑战。现有视频蒸馏(VD)方法往往计算开销大,难以保留时序动态,因直接套用图像方法常导致性能下降。本文提出一种新型单层级视频数据蒸馏框架,直接针对预训练模型优化合成视频。为缓解时间冗余并增强运动保留,引入基于帧间差异的时间显著性引导机制,鼓励保留关键时序信号同时抑制帧级冗余。在标准视频基准上的大量实验表明,该方法达到当前最优性能,缩小了真实与蒸馏视频数据间的差距,提供了一种可扩展的视频数据压缩解决方案。

原文摘要 · Abstract (English)

Dataset distillation (DD) has emerged as a powerful paradigm for dataset compression, enabling the synthesis of compact surrogate datasets that approximate the training utility of large-scale ones. While significant progress has been achieved in distilling image datasets, extending DD to the video domain remains challenging due to the high dimensionality and temporal complexity inherent in video data. Existing video distillation (VD) methods often suffer from excessive computational costs and struggle to preserve temporal dynamics, as naïve extensions of image-based approaches typically lead to degraded performance. In this paper, we propose a novel uni-level video dataset distillation framework that directly optimizes synthetic videos with respect to a pre-trained model. To address temporal redundancy and enhance motion preservation, we introduce a temporal saliency-guided filtering mechanism that leverages inter-frame differences to guide the distillation process, encouraging the retention of informative temporal cues while suppressing frame-level redundancy. Extensive experiments on standard video benchmarks demonstrate that our method achieves state-of-the-art performance, bridging the gap between real and distilled video data and offering a scalable solution for video dataset compression.

视频蒸馏数据压缩时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。