针对VMAF NEG在图像编码优化中的漏洞,提出更鲁棒的损失函数。
On Optimizing Image Codecs for VMAF NEG: Analysis, Issues, and a Robust Loss Proposal
- 发现VMAF NEG仍可被特定攻击误导,影响编码器训练效果。
- 提出融合VMAF NEG的鲁棒损失函数,提升编码质量感知一致性。
- 适用于追求高主观质量的图像编码器优化,尤其适合机器学习编码场景。
VMAF(视频多方法评估融合)指标因其与人类感知高度相关而日益受到图像和视频编码领域的关注,使其成为机器学习编码器训练与微调的理想目标。然而,研究发现该指标存在可被攻击的缺陷:例如,对图像进行去锐化处理虽能提升VMAF得分,却会降低真实感知质量。为此,研究人员提出了改进版的VMAF NEG,旨在增强对这类攻击的抵抗力,从而更适合作为编码器微调的评估标准。本文贡献有三:第一,系统分析了当前VMAF NEG在对抗性攻击下的潜在脆弱性,特别是当其被用于图像编码器微调时所引发的问题;第二,为充分利用VMAF NEG与人类感知的高度相关性,提出一种包含VMAF NEG的鲁棒损失函数,可用于编码器或解码器的微调;第三,通过多个图像样例的主观评价验证了量化结果的有效性,展示出该方法在提升感知质量方面的优势。
原文摘要 · Abstract (English)
The VMAF (video multi-method assessment fusion) metric for image and video coding recently gained more and more popularity as it is supposed to have a high correlation with human perception. This makes training and particularly fine-tuning of machine-learned codecs on this metric interesting. However, VMAF is shown to be attackable in a way that, e.g., unsharpening an image can lead to a gain in VMAF quality while decreasing the quality in human perception. A particular version of VMAF called VMAF NEG has been designed to be more robust against such attacks and therefore it should be more useful for fine-tuning of codecs. In this paper, our contributions are threefold. First, we identify and analyze the still existing vulnerability of VMAF NEG towards attacks, particulary towards the attack that consists in employing VMAF NEG for image codec fine-tuning. Second, to benefit from VMAF NEG's high correlation with human perception, we propose a robust loss including VMAF NEG for fine-tuning either the encoder or the decoder. Third, we support our quantitative objective results by providing perceptive impressions of some image examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。