用视觉语言强化学习评估视频编码质量,更贴近人眼判断。
CodecArena: Codec Quality Assessment via Visual Reinforcement Learning

- 将编码质量评估转化为参考帧与重建帧的对比推理任务。
- 在多种码率下对齐人类偏好,超越传统感知指标表现。
- 自动构建细粒度评价体系,适合研究视频生成与编码的学者。
视频编码正向低码率乃至超低码率发展,由端到端编码器取代手工流程,并利用生成模型先验的生成式编码器成为主流。然而主流评价指标如LPIPS和DISTS仅衡量特征与纹理相似性,无法反映内容保真度:即使重构出错误人脸或模糊文字为逼真笔触,仍可能获得高分,但人眼会立即拒绝。为此,我们提出CodecArena,首个用于视频编码质量评估的视觉-语言框架,将编码器评估建模为基于源帧的对比推理任务。通过Facet-GRPO这一视觉强化学习方案优化,对成对编码器偏好进行对齐,并以身份、物体、文本、纹理和时序一致性五个保真维度为基准。其基于自动推导的维度方向的奖励机制,避免单一分项得分主导整体评价,实现可解释的细粒度质量判断。为支持该未充分探索领域的训练与评估,我们构建了两个互补资源:CodecArena-1K,一个包含1500组比较的全自动偏好数据集,融合视觉-语言与客观监督;CodecArena-Bench,一个由源不重叠视频构成的人工排序基准,支持公平的域外评估。大量实验表明,CodecArena在跨源内容上与人类判断达成最优一致性,优于感知指标与先前视觉-语言评估器。
原文摘要 · Abstract (English)
Video coding is advancing into the low and ultra-low bitrate regime, driven by end-to-end codecs that replace the hand-crafted pipeline with jointly optimized neural networks and generative codecs that exploit the priors of video generation models. Yet the dominant metrics, LPIPS and DISTS, measure feature and texture similarity rather than content fidelity: a reconstruction that hallucinates a wrong face or blurs text into convincing strokes can still score well, even when a human rejects it instantly. To address this, we propose CodecArena, the first vision-language framework for video coding quality assessment, casting codec evaluation as source-conditioned comparative reasoning between a reference and its reconstructions. We optimize CodecArena with Facet-GRPO, a visual reinforcement learning scheme that aligns pairwise codec preferences while grounding the verdict in five fidelity facets: identity, objects, text, texture, and temporal consistency. Its facet-anchored reward uses automatically derived facet directions as weak anchors, rather than human per-facet labels, to prevent any single sub-score from dominating the holistic preference and to yield interpretable fine-grained quality judgments. To support training and evaluation in this underexplored regime, we construct two complementary resources: CodecArena-1K, a fully automatic preference dataset of 1,500 comparison groups built from traditional, neural, and generative codec reconstructions with fused vision-language and objective supervision; and CodecArena-Bench, a human-ranked benchmark with source-disjoint videos for fair out-of-domain evaluation. Extensive experiments demonstrate that CodecArena achieves state-of-the-art agreement with human judgments on source-disjoint content across diverse codecs and bitrates, surpassing perceptual metrics and prior vision-language evaluators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。