arXiv:2503.14112cs.CV2025-03CVPR被引 4

用生成模型压缩动作分割数据集,存储减500倍仍保83%性能

Condensing Action Segmentation Datasets via Generative Network Inversion

  • 用生成先验与网络反演将视频压缩为紧凑隐码
  • 在Breakfast数据集上存储减少500倍,性能保留83%
  • 适合需要高效训练的增量学习场景

本文提出首个用于时序动作分割(TAS)数据集的压缩方法。通过利用数据集学习到的生成先验和网络反演技术,将视频数据压缩为紧凑的隐码,在时间与通道维度均显著降低存储需求。同时,采用多样且具代表性的动作序列采样策略,有效减少视频层面的冗余。在标准基准上的评估表明,该方法在压缩TAS数据集方面表现稳定,并实现竞争性性能。具体而言,在Breakfast数据集上,存储量减少超过500倍,且训练性能保持在全数据集训练的83%。此外,在下游增量学习任务中,其表现优于当前最先进方法。

原文摘要 · Abstract (English)

This work presents the first condensation approach for procedural video datasets used in temporal action segmentation. We propose a condensation framework that leverages generative prior learned from the dataset and network inversion to condense data into compact latent codes with significant storage reduced across temporal and channel aspects. Orthogonally, we propose sampling diverse and representative action sequences to minimize video-wise redundancy. Our evaluation on standard benchmarks demonstrates consistent effectiveness in condensing TAS datasets and achieving competitive performances. Specifically, on the Breakfast dataset, our approach reduces storage by over 500$\times$ while retaining 83% of the performance compared to training with the full dataset. Furthermore, when applied to a downstream incremental learning task, it yields superior performance compared to the state-of-the-art.

数据压缩动作分割生成模型增量学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。