用大模型精准检测用户生成图像中的视觉失真,解决真实场景泛化难题。
Visual Distortion Detection in UGC Images Using Large Multimodal Models

- 利用多层大语言模型解码器作为多个检测器,同步分析多级特征。
- 构建14万张高质量失真图像数据集,覆盖8类主要合成失真类型。
- 保留非失真类别预测中的失真线索,缓解真实与合成图像间的差异问题。
局部感知质量评估在图像质量评估中长期面临关键挑战,现有基于大模型的方法多依赖文本驱动的监督微调,存在检测精度不足的问题。此外,常用合成失真图像训练数据在真实场景部署时表现出显著泛化差距,即“合成到真实(S2A)”难题。为此,我们提出VIGIL,利用大模型架构实现精确视觉失真检测。从超过100万样本池中构建了包含14万张失真图像的VIGIL-140K训练集,通过严格质量筛选和精心设计的失真注入,覆盖8类主要合成失真。模型利用大语言模型解码器不同层级作为多重检测器,基于多级特征同步进行失真检测;同时保留归为非失真类别的预测中隐含的失真线索,有效缓解S2A问题中常见的前景-背景分离模糊性。经过后处理,模型在合成失真检测和S2A任务上均持续优于强基线。
原文摘要 · Abstract (English)
The localized depiction of perceptual quality has long been a crucial, yet underexplored, challenge in image quality assessment (IQA). Existing approaches based on large multimodal models (LMMs) predominantly rely on text-driven supervised fine-tuning (SFT). However, this training paradigm exhibits notable limitations in detection accuracy. Moreover, synthetically distorted images, which are often used as the primary training data source, show a significant generalization gap when deployed in real-world scenarios; thus, the \textbf{synthetic-to-authentic (\textit{S2A})} problem represents a critical challenge. Motivated by these issues, we propose \textbf{\textit{VIGIL}}, which leverages the LMM architecture for precise visual distortion detection. From a candidate pool of over 1000K samples, we construct the \textbf{\textit{VIGIL-140K}} training set, which consists of over 140K distorted images. These images are obtained through rigorous quality filtering and carefully crafted distortion injection, covering 8 major synthetic distortion categories. Our model leverages different layers of the large language model (LLM) decoder, treating them as \textit{multiple detectors} that perform synchronous distortion detection using multi-level features. Additionally, we retain distortion cues from predictions assigned to the non-distortion class, which helps mitigate the ambiguous foreground-background (\textit{FG-BG}) separation commonly encountered in the \textit{S2A} problem. After post-processing, our model consistently outperforms strong baselines on both in-domain synthetic distortion detection and \textit{S2A} tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。