arXiv:2604.24123cs.CV2026-04

提出跨传统与神经视频编码的通用画质评估方法,适应HDR和SDR。

FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs

论文配图:FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs
图 1 · 摘自论文原文
  • 融合深度与手工特征,捕捉从纹理到语义的多尺度失真。
  • 在10个新编码器数据集上与主观评分相关性超0.95,泛化性强。
  • 适合视频编解码研究者,尤其关注神经编码器评估场景。

视频技术正向超高清(UHD)和高动态范围(HDR)发展,对高效率压缩的需求日益增强。除传统编码器外,神经视频编码器(NVCs)近年受到广泛关注并快速演进。NVC的编码伪影具有内容依赖性和生成特性,与传统编码器不同,难以被传统视频质量评估(VQA)方法准确捕捉。因此,亟需一种能跨编码器、内容类型和动态范围的通用VQA度量方法,以更好支持视频编码研究与评估。本文提出FDIM,一种基于特征距离的通用视频质量度量,适用于传统与神经视频编码器,覆盖SDR与HDR格式。FDIM采用混合架构,结合深度特征与手工特征:深度部分学习多尺度表示,捕捉从结构纹理退化到高层语义偏差的失真;手工特征部分提供稳定互补线索,提升整体泛化能力。我们在包含超过16,000段视频序列的大规模主观质量评估数据集DCVQA上训练FDIM,该数据集涵盖传统块基混合编码器与端到端感知优化的神经视频编码器。在10个包含多样、未见过编码器的SDR/HDR VQA数据集上进行大量实验,结果表明FDIM具备强大泛化能力,与主观评估相关性超过0.95。FDIM源代码及DCVQA验证集将发布于https://github.com/MCL-ZJU/FDIM。

原文摘要 · Abstract (English)

Video technology is advancing toward Ultra High Definition (UHD) and High Dynamic Range (HDR), which intensifies the need for higher compression efficiency for these high-specification videos. Beyond advances in traditional codecs, neural video codecs (NVCs) have attracted significant research attention and have evolved rapidly over the past few years. The coding artifacts of NVCs often exhibit content-varying and generative characteristics, which differ from those of conventional codecs and are challenging for traditional video quality assessment (VQA) methods to capture. Therefore, VQA metrics are required to generalize across different codecs, content types, and dynamic ranges to better support video codec research and evaluation. In this paper, we propose FDIM, a feature-distance-based generic video quality metric for both traditional and neural video codecs across SDR and HDR formats. FDIM employs a hybrid architecture that integrates deep and hand-crafted features. The deep feature component learns multi-scale representations to capture distortions ranging from structural and textural fidelity degradation to high-level semantic deviations, while the hand-crafted feature component provides stable complementary cues to improve overall generalization. We trained FDIM on a large-scale subjective quality assessment dataset (DCVQA) consisting of over 16k video sequences encoded by traditional block-based hybrid video codecs and end-to-end perceptually optimized neural video codecs. Extensive experiments on ten SDR/HDR VQA datasets containing diverse, previously unseen codecs demonstrate that FDIM achieves strong generalization and high correlation with subjective assessment. The source code for FDIM and the DCVQA validation set will be released at https://github.com/MCL-ZJU/FDIM.

视频质量神经编码泛化评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。