用压缩响应自监督聚类视频编码复杂度,提升流媒体效率
A Self-Supervised Learning Framework for Video Encoding Complexity Clustering

- 通过压缩回声作为自监督信号,学习视频编码特性
- 在多个数据集上优于现有视觉编码器,节省带宽与提升画质
- 适合视频流媒体系统优化编码策略的工程师
自适应视频流是互联网视频传输的常用技术。其关键挑战在于为每段视频确定最优编码设置,而编码复杂度因内容特征差异显著。本文提出压缩回声对比学习(CECL),一种基于视频压缩响应的自监督学习框架,用于按编码复杂度聚类视频。该方法利用视频对压缩的响应——压缩回声——作为监督信号,在预训练中捕捉深层编码特性。大量实验表明,所学表征在下游编码复杂度聚类任务中表现优异,显著优于现有先进视觉编码器,并在固定码率阶梯方案下实现显著码率与画质优势。
原文摘要 · Abstract (English)
Adaptive video streaming is a widely used technique for delivering video content over the internet. One of the key challenges is determining the optimal encoding settings for each video, which can vary significantly based on its content and characteristics. In this paper, we propose Compression Echo Contrastive Learning (CECL), a novel self-supervised learning framework for clustering videos based on their encoding complexity. Our method leverages the response of a video to compression - the Compression Echo - as a supervisory signal, allowing the model to capture underlying encoding characteristics during pretraining. We conduct extensive experiments to demonstrate the effectiveness of our learned representations for the downstream task of clustering videos by their encoding complexity. Our results show that CECL improves upon existing state-of-the-art visual encoders and delivers strong bitrate and quality savings against the fixed bitrate ladder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。