arXiv:2502.19026eess.IVcs.AI2025-02中稿 · ISCAS 2025被引 5

用大模型蒸馏出轻量级视频质量评估模型,提升压缩画质判断能力。

InternVQA: Advancing Compressed Video Quality Assessment with Distilling Large Foundation Model

  • 用InternVideo2大模型蒸馏出轻量级模型,保留压缩画质先验知识。
  • 在压缩视频质量评估任务上超越现有方法,性能显著提升。
  • 适合需要高效部署的视频质量检测场景,如流媒体平台。

视频质量评估依赖于丰富的视频理解特征,如语义信息、纹理和时序运动。现有的视频基础模型InternVideo2因其参数量大及基于大规模多模态数据训练,展现出强大的视频理解能力。本文探索了InternVideo2在压缩场景下视频质量评估任务中的迁移潜力。为构建适配该任务的轻量级模型,提出一种蒸馏方法,使小模型具备丰富的压缩质量先验知识。同时研究了不同骨干网络在蒸馏过程中的表现。结果表明,相较于其他方法,从InternVideo2蒸馏得到的轻量级模型在压缩视频质量评估任务中表现出色,性能优异。

原文摘要 · Abstract (English)

Video quality assessment tasks rely heavily on the rich features required for video understanding, such as semantic information, texture, and temporal motion. The existing video foundational model, InternVideo2, has demonstrated strong potential in video understanding tasks due to its large parameter size and large-scale multimodal data pertaining. Building on this, we explored the transferability of InternVideo2 to video quality assessment under compression scenarios. To design a lightweight model suitable for this task, we proposed a distillation method to equip the smaller model with rich compression quality priors. Additionally, we examined the performance of different backbones during the distillation process. The results showed that, compared to other methods, our lightweight model distilled from InternVideo2 achieved excellent performance in compression video quality assessment.

视频质量评估模型蒸馏轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。