arXiv:2603.25566eess.IV2026-03

用Mamba模型设计新感知损失,让视频转码更贴近人眼感受。

A Mamba-based Perceptual Loss Function for Learning-based UGC Transcoding

  • 基于Mamba的轻量级质量模型,弱监督训练提升感知能力。
  • 在DCVC和HiNeRV上实现8.46%与12.89%的BD-rate降低。
  • 适合关注视频转码主观画质优化的研究者或工程师。

在用户生成内容(UGC)转码中,源视频常因前期压缩、剪辑或拍摄条件不佳而存在多种退化。现有仅以参考视频为像素保真目标的压缩范式因此表现欠佳,因其迫使编码器复现源视频的固有伪影。为此,本文提出一种新的感知启发式损失函数,重新定义参考视频的角色:从真实像素锚点转变为信息性上下文引导。具体地,我们基于选择性结构状态空间模型(Mamba)训练一个轻量级神经质量模型,并采用弱监督孪生排序策略进行优化。该模型被集成至两种神经视频编码器(DCVC与HiNeRV)的率失真优化(RDO)流程中作为损失函数,旨在生成更具感知质量的重建内容。实验表明,该框架在自动编码器与隐式神经表示基线之上分别取得8.46%与12.89%的BD-rate节省,显著提升编码效率。

原文摘要 · Abstract (English)

In user-generated content (UGC) transcoding, source videos typically suffer various degradations due to prior compression, editing, or suboptimal capture conditions. Consequently, existing video compression paradigms that solely optimize for fidelity relative to the reference become suboptimal, as they force the codec to replicate the inherent artifacts of the non-pristine source. To address this, we propose a novel perceptually inspired loss function for learning-based UGC video transcoding that redefines the role of the reference video, shifting it from a ground-truth pixel anchor to an informative contextual guide. Specifically, we train a lightweight neural quality model based on a Selective Structured State-Space Model (Mamba) optimized using a weakly-supervised Siamese ranking strategy. The proposed model is then integrated into the rate-distortion optimization (RDO) process of two neural video codecs (DCVC and HiNeRV) as a loss function, aiming to generate reconstructed content with improved perceptual quality. Our experiments demonstrate that this framework achieves substantial coding gains over both autoencoder and implicit neural representation-based baselines, with 8.46% and 12.89% BD-rate savings, respectively.

视频转码感知损失Mamba率失真优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。