构建首个UGC图像失真分析数据集,实现细节化质量评估与解释。
ViDA-UGC: Detailed Image Quality Analysis via Visual Distortion Assessment for UGC Images
- 基于失真导向流程,结合人类标注与GPT-4o的思维链框架生成细粒度质量描述。
- 建立11,000张图像的ViDA-UGC数据集,含6,149个问答对,支持高质量分析。
- 适用于图像质量控制、修复指导及多模态大模型评测,适合视觉质量研究者。
多模态大语言模型(MLLMs)推动图像质量评估(IQA)从不可解释评分转向可解释评估,在质量控制与优化指导中展现应用潜力。然而,现有可解释IQA方法对用户生成内容(UGC)与AI生成内容(AIGC)使用相同失真标准,且缺乏对图像质量的细致分析以支持修复与监控。本文构建首个针对UGC图像的视觉失真评估指令微调数据集——ViDA-UGC,包含11,000张图像,涵盖细粒度质量标注、详细感知描述及推理性质量说明。该数据集通过失真导向流程构建,融合人工标注与思维链(CoT)评估框架,引导GPT-4o识别并分析UGC失真,捕捉与失真模式相关联的丰富低层视觉特征。同时,从数据集中精选476张图像,生成6,149个问答对,并由专业团队校验,形成首个UGC失真评估基准——ViDA-UGC-Bench。实验表明,该数据集与CoT框架能持续提升多种基础MLLM在ViDA-UGC-Bench和Q-Bench上的质量分析能力,甚至超越GPT-4o。
原文摘要 · Abstract (English)
Recent advances in Multimodal Large Language Models (MLLMs) have introduced a paradigm shift for Image Quality Assessment (IQA) from unexplainable image quality scoring to explainable IQA, demonstrating practical applications like quality control and optimization guidance. However, current explainable IQA methods not only inadequately use the same distortion criteria to evaluate both User-Generated Content (UGC) and AI-Generated Content (AIGC) images, but also lack detailed quality analysis for monitoring image quality and guiding image restoration. In this study, we establish the first large-scale Visual Distortion Assessment Instruction Tuning Dataset for UGC images, termed ViDA-UGC, which comprises 11K images with fine-grained quality grounding, detailed quality perception, and reasoning quality description data. This dataset is constructed through a distortion-oriented pipeline, which involves human subject annotation and a Chain-of-Thought (CoT) assessment framework. This framework guides GPT-4o to generate quality descriptions by identifying and analyzing UGC distortions, which helps capturing rich low-level visual features that inherently correlate with distortion patterns. Moreover, we carefully select 476 images with corresponding 6,149 question answer pairs from ViDA-UGC and invite a professional team to ensure the accuracy and quality of GPT-generated information. The selected and revised data further contribute to the first UGC distortion assessment benchmark, termed ViDA-UGC-Bench. Experimental results demonstrate the effectiveness of the ViDA-UGC and CoT framework for consistently enhancing various image quality analysis abilities across multiple base MLLMs on ViDA-UGC-Bench and Q-Bench, even surpassing GPT-4o.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。