提出视觉编码器为中心的预训练方法,提升视觉质量评估模型泛化能力。
VITAL: Vision-Encoder-centered Pre-training for LMMs in Visual Quality Assessment
- 以视觉编码器为核心,构建超大规模图文对数据集。
- 多任务训练使模型在图像与视频上兼具评分精度与解释能力。
- 仅需极少数据微调即可达到接近全量训练效果,适合快速部署。
构建稳健的视觉质量评估(VQualA)大模型需要兼顾多样性、强大性与可迁移性。然而现有VQualA大模型多聚焦单一任务,依赖全参数微调,易在特定模态或任务上过拟合,限制泛化与迁移能力。为此,我们提出以视觉编码器为中心的生成式预训练流程,开发VITAL系列大模型:(1)采用机器执行的标注审查范式,构建超过450万条视觉-语言配对数据,是迄今最大的VQualA训练数据集;(2)采用多任务训练策略,同步提升模型在图像与视频上的量化评分精度和质量解释能力;(3)基于视觉编码器实现高效模型库扩展:模型库具备强零样本性能,每个配对解码器仅需少于千分之一预训练数据的快速冷启动,即可达到全量训练模型水平。整体工作为构建视觉质量评估基础大模型奠定关键基础。
原文摘要 · Abstract (English)
Developing a robust visual quality assessment (VQualA) large multi-modal model (LMM) requires achieving versatility, powerfulness, and transferability. However, existing VQualA LMMs typically focus on a single task and rely on full-parameter fine-tuning, which makes them prone to overfitting on specific modalities or task types, thereby limiting their generalization capacity and transferability. To address this, we propose a vision-encoder-centered generative pre-training pipeline and develop the VITAL-Series LMMs. (1) We adopt a machine-executed annotation-scrutiny paradigm, constructing over 4.5M vision-language (VL) pairs-the largest VQualA training dataset to date. (2) We employ a multi-task training workflow that simultaneously enhances the model's quantitative scoring precision and strengthens its capability for quality interpretation across both image and video modalities. (3) Building upon the vision encoder, we realize an efficient model zoo extension: the model zoo exhibits strong zero-shot performance, and each paired decoder requires only a swift warm-up using less than 1/1000 of the pre-training data to achieve performance comparable to the fully trained counterpart. Overall, our work lays a cornerstone for advancing toward the foundation LMM for VQualA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。