用扩散模型生成视频的隐式函数表示,实现超低码率压缩。
Compression as Adaptation: Implicit Visual Representation with Diffusion Foundation Models
- 将图像/视频编码为冻结生成模型的低秩适配函数。
- 81帧视频可压缩为单个紧凑向量,实现极低码率下的保真压缩。
- 支持推理时调整,适合需要灵活压缩与生成的场景。
现代视觉生成模型通过大规模训练获得丰富的视觉知识,但现有的视觉表示(如像素、潜在变量或标记)仍位于模型外部,无法直接利用这些知识实现紧凑存储或复用。本文提出一种新的视觉表示框架,将信号编码为函数形式,该函数由附加在冻结视觉生成模型上的低秩适配参数化。这种隐式表示(例如81帧视频)可进一步哈希为单一紧凑向量,在极低比特率下实现优异的感知视频压缩效果。除基础压缩外,该表示的函数特性支持推理时的动态扩展与控制,可进一步优化压缩性能。更广泛地,由于隐式表示直接作用于生成过程,这为视觉压缩与生成提供了一个统一框架。
原文摘要 · Abstract (English)
Modern visual generative models acquire rich visual knowledge through large-scale training, yet existing visual representations (such as pixels, latents, or tokens) remain external to the model and cannot directly exploit this knowledge for compact storage or reuse. In this work, we introduce a new visual representation framework that encodes a signal as a function, which is parametrized by low-rank adaptations attached to a frozen visual generative model. Such implicit representations of visual signals, \textit{e.g.}, an 81-frame video, can further be hashed into a single compact vector, achieving strong perceptual video compression at extremely low bitrates. Beyond basic compression, the functional nature of this representation enables inference-time scaling and control, allowing additional refinement on the compression performance. More broadly, as the implicit representations directly act as a function of the generation process, this suggests a unified framework bridging visual compression and generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。