arXiv:2510.12379eess.IVcs.AI2025-10中稿 · PCS 2025 Camera-Re…

轻量模型精准预测视频编码参数,实现高质量低功耗传输

LiteVPNet: A Lightweight Network for Video Encoding Control in Quality-Critical Applications

  • 用比特流特征、视频复杂度和语义嵌入设计轻量网络
  • 预测误差低于1.2分VMAF,87%测试样本误差在2分内
  • 适合影视级内容高效传输,尤其看重画质与能效的场景

近十年来,电影制作生态中的视频工作流程催生了视频流媒体的新应用场景,如现场虚拟制作。这类场景要求精确的质量控制与能效平衡。现有转码方法常因缺乏质量控制或计算开销过大而难以满足需求。为此,我们提出轻量级神经网络LiteVPNet,用于精准预测NVENC AV1编码器的量化参数,以达到指定VMAF评分。该模型采用低复杂度特征,包括比特流特性、视频复杂度度量及基于CLIP的语义嵌入。实验表明,LiteVPNet在多种质量目标下平均VMAF误差低于1.2分;在超过87%的测试数据中,误差控制在2分以内,优于当前最优方法(约61%)。其在不同质量区间均表现稳定,适用于高价值内容的高效、节能传输与流媒体体验提升。

原文摘要 · Abstract (English)

In the last decade, video workflows in the cinema production ecosystem have presented new use cases for video streaming technology. These new workflows, e.g. in On-set Virtual Production, present the challenge of requiring precise quality control and energy efficiency. Existing approaches to transcoding often fall short of these requirements, either due to a lack of quality control or computational overhead. To fill this gap, we present a lightweight neural network (LiteVPNet) for accurately predicting Quantisation Parameters for NVENC AV1 encoders that achieve a specified VMAF score. We use low-complexity features, including bitstream characteristics, video complexity measures, and CLIP-based semantic embeddings. Our results demonstrate that LiteVPNet achieves mean VMAF errors below 1.2 points across a wide range of quality targets. Notably, LiteVPNet achieves VMAF errors within 2 points for over 87% of our test corpus, c.f. approx 61% with state-of-the-art methods. LiteVPNet's performance across various quality regions highlights its applicability for enhancing high-value content transport and streaming for more energy-efficient, high-quality media experiences.

视频编码轻量模型AV1质量控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。