arXiv:2409.11718cs.CV2024-09ECCV被引 18

利用视觉基础模型提升无监督视频语义压缩的语义丰富度

Free-VSC: Free Semantics from Visual Foundation Models for Unsupervised Video Semantic Compression

  • 通过共享对齐层与特定提示,融合多个视觉基础模型语义
  • 在六个数据集上超越现有方法,降低编码比特率
  • 适合需要高效视频压缩与多任务分析的场景

无监督视频语义压缩(UVSC)近年来受到关注,但以往方法语义能力有限,受限于单一语义目标和训练数据不足。为此,我们提出利用现成的视觉基础模型(VFMs)丰富的语义信息来增强UVSC任务。具体地,引入一个跨模型共享的语义对齐层,并搭配各VFM特有的提示,灵活对齐压缩视频与不同VFMs之间的语义。这使得多个VFMs协同构建相互增强的语义空间,指导压缩模型学习。此外,设计了一种基于动态轨迹的帧间压缩方案:先根据历史内容估计语义轨迹,再沿轨迹预测未来语义作为编码上下文,从而降低系统整体比特开销,进一步提升压缩效率。所提方法在三个主流任务和六个数据集上均优于现有编码方法。

原文摘要 · Abstract (English)

Unsupervised video semantic compression (UVSC), i.e., compressing videos to better support various analysis tasks, has recently garnered attention. However, the semantic richness of previous methods remains limited, due to the single semantic learning objective, limited training data, etc. To address this, we propose to boost the UVSC task by absorbing the off-the-shelf rich semantics from VFMs. Specifically, we introduce a VFMs-shared semantic alignment layer, complemented by VFM-specific prompts, to flexibly align semantics between the compressed video and various VFMs. This allows different VFMs to collaboratively build a mutually-enhanced semantic space, guiding the learning of the compression model. Moreover, we introduce a dynamic trajectory-based inter-frame compression scheme, which first estimates the semantic trajectory based on the historical content, and then traverses along the trajectory to predict the future semantics as the coding context. This reduces the overall bitcost of the system, further improving the compression efficiency. Our approach outperforms previous coding methods on three mainstream tasks and six datasets.

视频压缩语义压缩视觉模型无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。