通过对抗性权重扰动提升视频质量评估的跨数据集泛化能力
Revisiting Video Quality Assessment from the Perspective of Generalization
- 从泛化能力视角重看VQA,分析模型权重损失景观
- 对抗性权重扰动使损失景观更平滑,跨数据集性能提升1.8%
- 方法通用性强,适用于多种VQA/IQA模型,适合工业级应用
短视频平台如YouTube Shorts、TikTok和Kwai的兴起带来了大量用户生成内容(UGC),对视频质量评估(VQA)任务的泛化性能提出了严峻挑战。现有研究多关注特征提取器、采样策略和网络分支设计,却忽视了泛化能力本身。本文从泛化角度重新审视VQA任务,首先分析模型权重损失景观,发现其与泛化差距存在强相关性;随后探索多种正则化方法以平滑该景观。实验表明,对抗性权重扰动可有效优化损失景观,显著提升泛化性能:跨数据集测试性能最高提升1.8%,微调性能最高提升3%。在多种VQA方法与数据集上的广泛验证证明了方法的有效性。进一步地,基于此洞察,我们在图像质量评估(IQA)任务中也取得了当前最优结果。
原文摘要 · Abstract (English)
The increasing popularity of short video platforms such as YouTube Shorts, TikTok, and Kwai has led to a surge in User-Generated Content (UGC), which presents significant challenges for the generalization performance of Video Quality Assessment (VQA) tasks. These challenges not only affect performance on test sets but also impact the ability to generalize across different datasets. While prior research has primarily focused on enhancing feature extractors, sampling methods, and network branches, it has largely overlooked the generalization capabilities of VQA tasks. In this work, we reevaluate the VQA task from a generalization standpoint. We begin by analyzing the weight loss landscape of VQA models, identifying a strong correlation between this landscape and the generalization gaps. We then investigate various techniques to regularize the weight loss landscape. Our results reveal that adversarial weight perturbations can effectively smooth this landscape, significantly improving the generalization performance, with cross-dataset generalization and fine-tuning performance enhanced by up to 1.8% and 3%, respectively. Through extensive experiments across various VQA methods and datasets, we validate the effectiveness of our approach. Furthermore, by leveraging our insights, we achieve state-of-the-art performance in Image Quality Assessment (IQA) tasks. Our code is available at https://github.com/XinliYue/VQA-Generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。