用信息独特性提升视频压缩效率,更省算力还保画质。
UniComp: Rethinking Video Compression Through Informational Uniqueness
- 基于信息独特性设计三模块,自适应分配压缩资源。
- 在有限算力下保留更多关键视觉信息,重建误差更低。
- 适合追求高效低延迟视频压缩的研究与工程人员。
与基于注意力的压缩方法不同,本文提出一种以信息独特性驱动的视频压缩框架UniComp,旨在有限计算预算下最大化视频表征的信息保真度。从信息论角度出发,将视频压缩建模为最小化保留令牌与完整令牌间条件熵(重建误差)的优化问题。为此,引入信息独特性概念来度量令牌间的内在冗余,并关联到重建误差。基于此,设计了帧组融合、令牌分配和空间动态压缩三个模块,依次实现语义帧分组、自适应资源分配和细粒度空间压缩。大量实验表明,UniComp在有限计算预算下持续优于现有压缩方法,能更好地保留关键视觉令牌,凸显信息独特性在令牌压缩效能中的核心作用。
原文摘要 · Abstract (English)
Distinct from attention-based compression methods, this paper presents an information uniqueness driven video compression framework, termed UniComp, which aims to maximize the information fidelity of video representations under constrained computational budgets. Starting from the information-theoretic perspective, we formulate the vision compression as an optimization problem that minimizes conditional entropy (reconstruction error) between retained and full tokens. To achieve this, we introduce the notion of information uniqueness to measure intrinsic redundancy among tokens to link with reconstruction error. Based on uniqueness, we design three modules-Frame Group Fusion, Token Allocation, and Spatial Dynamic Compression-that progressively perform semantic frame grouping, adaptive resource allocation, and fine-grained spatial compression. Extensive experiments demonstrate that UniComp consistently outperforms existing compression methods in preserving essential visual tokens under limited computational budgets, highlighting the pivotal role of information uniqueness in token compression efficacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。