arXiv:2411.09308eess.IVcs.CV2024-11被引 5

用深度变压器模型预测机器视觉可识别的最小差异,降低视频编码比特率。

DT-JRD: Deep Transformer based Just Recognizable Difference Prediction Model for Video Coding for Machines

  • 基于多分类的深度变换器模型,融合内容与失真特征。
  • 预测误差仅5.574,比现有模型低13.1%;编码节省29.58%比特率。
  • 适合需要高效视频编码的机器视觉任务,如目标检测。

Just Recognizable Difference(JRD)表示机器视觉能检测到的最小视觉差异,可用于提升面向机器的视觉信号处理效率。本文提出基于深度变换器的JRD预测模型(DT-JRD),用于视频编码为机器(VCM)。首先,将JRD预测建模为多分类问题,设计集成改进嵌入、内容与失真特征提取、多分类及新型学习策略的模型。其次,受机器视觉在接近JRD时对失真响应相似的感知特性启发,提出基于高斯分布软标签(GDSL)的渐近JRD损失函数,显著扩展训练标签数量并放宽分类边界。最后,构建基于DT-JRD的VCM框架,在保持目标检测精度前提下降低编码比特率。大量实验表明,所提模型预测的平均绝对误差为5.574,优于当前最优模型13.1%;与VVC相比,平均比特率降低29.58%。

原文摘要 · Abstract (English)

Just Recognizable Difference (JRD) represents the minimum visual difference that is detectable by machine vision, which can be exploited to promote machine vision oriented visual signal processing. In this paper, we propose a Deep Transformer based JRD (DT-JRD) prediction model for Video Coding for Machines (VCM), where the accurately predicted JRD can be used reduce the coding bit rate while maintaining the accuracy of machine tasks. Firstly, we model the JRD prediction as a multi-class classification and propose a DT-JRD prediction model that integrates an improved embedding, a content and distortion feature extraction, a multi-class classification and a novel learning strategy. Secondly, inspired by the perception property that machine vision exhibits a similar response to distortions near JRD, we propose an asymptotic JRD loss by using Gaussian Distribution-based Soft Labels (GDSL), which significantly extends the number of training labels and relaxes classification boundaries. Finally, we propose a DT-JRD based VCM to reduce the coding bits while maintaining the accuracy of object detection. Extensive experimental results demonstrate that the mean absolute error of the predicted JRD by the DT-JRD is 5.574, outperforming the state-of-the-art JRD prediction model by 13.1%. Coding experiments shows that comparing with the VVC, the DT-JRD based VCM achieves an average of 29.58% bit rate reduction while maintaining the object detection accuracy.

视频编码机器视觉深度学习变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。