多任务机器视觉视频编码,提升精度与压缩效率。
Multi-task Just Recognizable Difference for Video Coding for Machines: Database, Model, and Coding Application
- 构建支持三任务的多任务JRD数据集,含27264条标注。
- 提出融合属性信息的联合预测模型,平均误差仅3.781。
- 适用于需要高效视频传输的机器视觉系统。
Just Recognizable Difference (JRD) 通过可见性阈值建模提升了机器视觉的编码效率,但当前仅限于单任务场景。为此,本文提出多任务JRD(MT-JRD)数据集与属性辅助的多任务JRD(AMT-JRD)模型,用于视频编码为机器(VCM),在提升预测精度的同时增强编码效率。首先,构建包含27,264条机器标注的MT-JRD数据集,涵盖目标检测、实例分割和关键点检测三项代表性任务。其次,提出AMT-JRD预测模型,融合通用特征提取模块(GFEM)与专用特征提取模块(SFEM),实现多任务联合学习。第三,引入属性特征融合模块(AFFM),将物体尺寸与位置等先验信息融入逐对象的JRD预测,弥补仅依赖图像特征的不足,增强对机器视觉感知机制的建模能力。最后,将AMT-JRD应用于VCM,利用精确预测的JRD降低码率,同时保持多任务精度。实验表明,该模型在三项任务上平均绝对误差为3.781,误差方差5.332,优于现有单任务模型6.7%和6.3%。编码测试显示,相比基准的VVC与JPEG,AMT-JRD-based VCM分别提升3.861%与7.886%的Bjontegaard Delta-mean Average Precision(BD-mAP)。
原文摘要 · Abstract (English)
Just Recognizable Difference (JRD) boosts coding efficiency for machine vision through visibility threshold modeling, but is currently limited to a single-task scenario. To address this issue, we propose a Multi-Task JRD (MT-JRD) dataset and an Attribute-assisted MT-JRD (AMT-JRD) model for Video Coding for Machines (VCM), enhancing both prediction accuracy and coding efficiency. First, we construct a dataset comprising 27,264 JRD annotations from machines, supporting three representative tasks including object detection, instance segmentation, and keypoint detection. Secondly, we propose the AMT-JRD prediction model, which integrates Generalized Feature Extraction Module (GFEM) and Specialized Feature Extraction Module (SFEM) to facilitate joint learning across multiple tasks. Thirdly, we innovatively incorporate object attribute information into object-wise JRD prediction through the Attribute Feature Fusion Module (AFFM), which introduces prior knowledge about object size and location. This design effectively compensates for the limitations of relying solely on image features and enhances the model's capacity to represent the perceptual mechanisms of machine vision. Finally, we apply the AMT-JRD model to VCM, where the accurately predicted JRDs are applied to reduce the coding bit rate while preserving accuracy across multiple machine vision tasks. Extensive experimental results demonstrate that AMT-JRD achieves precise and robust multi-task prediction with a mean absolute error of 3.781 and error variance of 5.332 across three tasks, outperforming the state-of-the-art single-task prediction model by 6.7% and 6.3%, respectively. Coding experiments further reveal that compared to the baseline VVC and JPEG, the AMT-JRD-based VCM improves an average of 3.861% and 7.886% Bjontegaard Delta-mean Average Precision (BD-mAP), respectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。