arXiv:2505.22564cs.CVcs.AI2025-05被引 2

PRISM通过动态插入关键帧,高效压缩视频数据。

PRISM: Video Dataset Condensation with Progressive Refinement and Insertion for Sparse Motion

  • 将视频视为时空整体,按需插入关键帧而非固定帧优化
  • 仅在非线性运动处插入帧,实现存储效率与精度平衡
  • 适合需要压缩视频数据的场景,如边缘计算、视频检索

视频数据压缩旨在降低视频处理的巨大计算成本。然而,其面临空间外观与时间动态之间不可分割的耦合挑战。以往方法采用静态/动态解耦范式,将视频分解为静态内容和辅助运动信号,这种多阶段方法常扭曲真实动作的内在关联。我们提出一种整体性方法——针对稀疏运动的渐进式精炼与插入(PRISM),从一开始就将视频视为统一且完全耦合的时空结构。为最大化表征效率,PRISM通过避免固定帧优化来应对视频固有的时间冗余。它从最少的时间锚点开始,仅在直线插值无法捕捉非线性动态时才逐步插入关键帧,关键时刻由梯度错位识别。该自适应过程确保表征能力精准分配至必要位置,最小化存储需求的同时保留复杂运动。大量实验表明,PRISM在标准基准上表现优异,同时通过稀疏且整体学习的表示实现最先进的存储效率。

原文摘要 · Abstract (English)

Video dataset condensation aims to reduce the immense computational cost of video processing. However, it faces a fundamental challenge regarding the inseparable interdependence between spatial appearance and temporal dynamics. Prior work follows a static/dynamic disentanglement paradigm where videos are decomposed into static content and auxiliary motion signals. This multi-stage approach often misrepresents the intrinsic coupling of real-world actions. We introduce Progressive Refinement and Insertion for Sparse Motion (PRISM), a holistic approach that treats the video as a unified and fully coupled spatiotemporal structure from the outset. To maximize representational efficiency, PRISM addresses the inherent temporal redundancy of video by avoiding fixed-frame optimization. It begins with minimal temporal anchors and progressively inserts key-frames only where linear interpolation fails to capture non-linear dynamics. These critical moments are identified through gradient misalignments. Such an adaptive process ensures that representational capacity is allocated precisely where needed, minimizing storage requirements while preserving complex motion. Extensive experiments demonstrate that PRISM achieves competitive performance across standard benchmarks while providing state-of-the-art storage efficiency through its sparse and holistically learned representation.

视频压缩时空建模数据凝缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。