用大模型自动优化视频编码的量化参数分配,提升压缩效率
LLM-Driven Heuristic Frame-Level Quantization Parameter Adaptation for VVenC

- 用大模型生成可执行的编码启发式算法,通过真实编码反馈迭代优化
- 在多个测试集上优于固定参数和传统拉格朗日方法,实现更优码率-失真平衡
- 发现熵相关惩罚项能有效减少参数波动,为编码器设计提供新思路
帧级量化参数(QP)分配在现代视频编码器中仍是一个持续挑战。广泛采用的固定QP方案本质上与内容无关,而经典的拉格朗日率失真优化(RDO)方法常因乘子设置不准导致性能受限。本文探索使用大语言模型(LLM)自动设计帧级QP自适应的RDO启发式方法。构建了一个闭环进化框架:LLM迭代提出带有可执行代码的算法构想,这些候选方案通过弗劳恩霍夫通用视频编码器(VVenC)进行实际编码评估,每个启发式作为评分函数,基于历史帧和当前候选的编码统计信息比较不同QP选择。实验结果表明,在多个测试集上,演化出的启发式方法在码率-失真性能上均优于固定QP方案和拉格朗日基线。进一步分析显示,LLM自主发现了通过基于熵的项惩罚QP波动的自适应启发式,为RDO算法设计提供了新洞见。
原文摘要 · Abstract (English)
Optimal frame-level quantization parameter (QP) allocation remains a persistent challenge in modern video encoders. The fixed-QP scheme widely adopted in practical systems is inherently content-agnostic, while classical Lagrangian rate-distortion optimization (RDO) methods often suffer from inaccurate multiplier settings. In this paper, we explore the use of large language models (LLMs) to automatically design RDO heuristics for frame-level QP adaptation. We construct a closed-loop evolutionary framework in which the LLM iteratively proposes RDO heuristics as algorithmic ideas with executable code, and these candidates are evaluated directly through encoding with the Fraunhofer Versatile Video Encoder (VVenC), where each heuristic acts as a scoring function that compares different QP choices based on the encoding statistics of past frames and current candidates. Experimental results across multiple test sets show that the evolved heuristic achieves promising rate-distortion improvements over both the fixed-QP scheme and the Lagrangian baseline. Further analysis reveals that the LLM can autonomously discover an adaptive heuristic that penalizes QP fluctuations via entropy-based terms, providing new insights into the design of RDO algorithms
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。