深度学习视频滤波如何在低功耗设备上高效部署。
Deep Learning-based Filtering for Video Coding: A Survey on Architectures, Algorithms, and Complexity Analysis

- 从编码集成、信息利用到网络设计,三维度分类DLF方法。
- 轻量级模型相比传统模型降低40%以上功耗,适配NPU硬件。
- 面向8K设备与下一代智能终端,提供实用部署指南。
随着超高清(UHD)显示和沉浸式媒体服务在物联网(IoT)与消费电子(CE)领域普及,包括8K显示屏和移动设备,高效视频编码需求空前增长。深度学习视频滤波(DLF)已成为缓解HEVC/H.265和VVC/H.266等标准压缩伪影的有力方案,但其在消费电子设备中的部署受限于计算复杂度、内存带宽和功耗。为弥合学术研究与实际应用之间的差距,本文提出一项系统性、面向硬件的DLF技术综述。构建三维分类体系:(1)编码集成方式,(2)编码信息利用,(3)网络设计策略。区别于以往综述,本文重点分析率失真(RD)性能与硬件可行性间的权衡,揭示从高算力性能模型向轻量级、适配神经处理单元(NPUs)架构的演进趋势。同时整合联合视频专家组(JVET)关于基于神经网络的视频编码(NNVC)的最新标准化进展,提供切实可行的指导。识别出实时推理延迟与误差传播等开放挑战,为下一代低功耗智能视频编码在消费电子终端的发展指明路径。
原文摘要 · Abstract (English)
As Ultra-High-Definition (UHD) displays and immersive media services become ubiquitous in the Internet of Things (IoT) and Consumer Electronics (CE) sectors, including 8K display and mobile devices, the demand for high-efficiency video coding is unprecedented. While Deep Learning-based Filtering (DLF) has emerged as a promising solution to mitigate compression artifacts inherent in standards like High Efficiency Video Coding (HEVC/H.265) and Versatile Video Coding (VVC/H.266), its deployment in CE devices is severely constrained by computational complexity, memory bandwidth, and power consumption. To bridge the gap between academic research and practical deployment, this paper presents a comprehensive, hardware-oriented survey of DLF techniques. We propose a systematic three-dimensional taxonomy classifying methods into (1) Integration Scheme within the Video Coding, (2) Coding Information Utilization, and (3) Network Design Strategy. Unlike prior reviews, this work critically analyzes the trade-offs between Rate-Distortion (RD) performance and hardware feasibility, highlighting the evolution from heavy, performance-oriented models to lightweight, hardware-friendly architectures targeting Neural Processing Units (NPUs). Furthermore, we incorporate the latest standardization activities from the Joint Video Experts Team (JVET) on Neural Network-based Video Coding (NNVC) to provide realistic guidelines. We also identify open challenges such as real-time inference latency and error propagation, providing a roadmap toward robust, low-power intelligent video coding in next-generation CE vision endpoints.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。