基于模型反馈筛选难样本,提升视频质量评估精度
MDS-VQA: Model-Informed Data Selection for Video Quality Assessment
- 用失败预测器识别模型难处理的视频片段
- 5%精选数据使相关指标提升,显著改善模型表现
- 适合需要高效标注的高质量视频评估研究者
基于学习的视频质量评估(VQA)发展迅速,但模型设计与数据集构建之间存在脱节。现有方法多在固定基准上迭代模型,或盲目收集新的人类标签,未系统针对现有模型弱点。本文提出MDS-VQA,一种模型驱动的数据选择机制,用于筛选对基础模型困难且内容多样化的未标注视频。通过排名目标训练的失败预测器估计难度,利用深度语义特征衡量多样性,以贪婪算法在有限标注预算下平衡二者。在多个VQA数据集和模型上的实验表明,MDS-VQA能有效识别出具有挑战性且多样化的样本,特别适用于主动微调。仅使用每领域5%的精选子集,微调后模型的均值SRCC从0.651提升至0.722,并获得最优gMAD排名,表明其具备强适应性与泛化能力。
原文摘要 · Abstract (English)
Learning-based video quality assessment (VQA) has advanced rapidly, yet progress is increasingly constrained by a disconnect between model design and dataset curation. Model-centric approaches often iterate on fixed benchmarks, while data-centric efforts collect new human labels without systematically targeting the weaknesses of existing VQA models. Here, we describe MDS-VQA, a model-informed data selection mechanism for curating unlabeled videos that are both difficult for the base VQA model and diverse in content. Difficulty is estimated by a failure predictor trained with a ranking objective, and diversity is measured using deep semantic video features, with a greedy procedure balancing the two under a constrained labeling budget. Experiments across multiple VQA datasets and models demonstrate that MDS-VQA identifies diverse, challenging samples that are particularly informative for active fine-tuning. With only a 5% selected subset per target domain, the fine-tuned model improves mean SRCC from 0.651 to 0.722 and achieves the top gMAD rank, indicating strong adaptation and generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。