AI驱动的视频压缩模型实现高效编码,推动视觉智能发展。
Emerging Advances in Learned Video Compression: Models, Systems and Beyond
- 用端到端神经网络优化视频压缩,支持单向与双向预测。
- 实验显示其压缩性能显著优于传统方法,支持标准化进展。
- 适合关注AI视频技术、系统部署与硬件实现的研究者。
视频压缩是视觉智能的基础,连接视觉信号采集与高层视觉分析。人工智能技术的广泛应用,推动视频压缩进入新范式,通过端到端优化的神经模型实现突破。本文系统综述了近年来基于端到端学习的视频编码研究,涵盖单向与双向预测架构的开创性工作。重点分析了学习型视频压缩(LVC)中的优化技术,强调其技术创新与优势,并报告了相关标准化进展。同时探讨了LVC在系统设计与硬件实现中的挑战。最后通过大量仿真结果,验证了LVC模型在压缩性能上的卓越表现,回答了为何学习型编解码器与基于AI的视频技术将对未来的视觉智能研究产生深远影响。
原文摘要 · Abstract (English)
Video compression is a fundamental topic in the visual intelligence, bridging visual signal sensing/capturing and high-level visual analytics. The broad success of artificial intelligence (AI) technology has enriched the horizon of video compression into novel paradigms by leveraging end-to-end optimized neural models. In this survey, we first provide a comprehensive and systematic overview of recent literature on end-to-end optimized learned video coding, covering the spectrum of pioneering efforts in both uni-directional and bi-directional prediction based compression model designation. We further delve into the optimization techniques employed in learned video compression (LVC), emphasizing their technical innovations, advantages. Some standardization progress is also reported. Furthermore, we investigate the system design and hardware implementation challenges of the LVC inclusively. Finally, we present the extensive simulation results to demonstrate the superior compression performance of LVC models, addressing the question that why learned codecs and AI-based video technology would have with broad impact on future visual intelligence research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。