用视频质量评估提升视频分类准确率,解决标注不足问题。
Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition
- 结合自监督学习与无参考画质评估,联合优化分类任务。
- 在I-CONECT数据集上达94.87%准确率,显著提升模糊视频识别效果。
- 适合医疗视频分析、弱标注视频分类等场景使用。
视频质量显著影响视频分类效果。我们发现,用清晰视频能有效识别轻度认知障碍,而模糊视频则表现较差。由此意识到引入视频质量评估(VQA)可提升分类性能。本文提出基于自监督学习的视频视觉变压器联合无参考VQA进行视频分类的方法(SSL-V3)。SSL-V3采用联合自监督机制,将视频质量评分作为调节因子直接作用于分类特征图,解决视频数据集中常见且难以获取标签的VQA标注缺失问题。该评分作为桥梁,使分类任务反向优化VQA参数。实验表明,该方法在两个数据集上均表现稳健,尤其在包含面部信息的医疗数据集I-CONECT中,对访谈视频的分类准确率达94.87%,验证了其有效性。
原文摘要 · Abstract (English)
Video quality significantly affects video classification. We found this problem when we classified Mild Cognitive Impairment well from clear videos, but worse from blurred ones. From then, we realized that referring to Video Quality Assessment (VQA) may improve video classification. This paper proposed Self-Supervised Learning-based Video Vision Transformer combined with No-reference VQA for video classification (SSL-V3) to fulfill the goal. SSL-V3 leverages Combined-SSL mechanism to join VQA into video classification and address the label shortage of VQA, which commonly occurs in video datasets, making it impossible to provide an accurate Video Quality Score. In brief, Combined-SSL takes video quality score as a factor to directly tune the feature map of the video classification. Then, the score, as an intersected point, links VQA and classification, using the supervised classification task to tune the parameters of VQA. SSL-V3 achieved robust experimental results on two datasets. For example, it reached an accuracy of 94.87% on some interview videos in the I-CONECT (a facial video-involved healthcare dataset), verifying SSL-V3's effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。