arXiv:2603.10578cs.CVcs.DB2026-03被引 1

用检索增强生成提升视觉语言模型对图像质量的细粒度判断能力

R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment

  • 基于用户感知构建6维质量维度,标注3500张渲染图
  • 检索相似图像描述可显著提升模型对细微质量差异的判断
  • 提出双流检索框架,适配高质量图像评估任务

沉浸式计算机图形渲染已广泛应用于现代生活。然而,全面评估其质量仍面临两大挑战:一是现有CG数据集缺乏系统性的渲染质量描述;二是现有质量评估方法无法提供合理的文本解释。为此,我们从用户视角识别出六项关键感知维度,构建了包含3500张CG图像及其对应质量描述的数据集,每条描述涵盖图像风格、内容及所选维度的感知质量。此外,利用该数据集子集构建多个基于描述的问答基准,用于评估现有视觉语言模型(VLMs)的表现。实验发现,当前VLMs在细粒度质量判断上仍不充分,但视觉相似图像的描述能显著提升其对目标图像的理解。受此启发,我们采用检索增强生成,提出一种双流检索框架,有效增强了VLMs的CG质量评估能力。在多个代表性VLM上的实验表明,该方法显著提升了其在CG质量评估任务中的表现。

原文摘要 · Abstract (English)

Immersive Computer Graphics (CGs) rendering has become ubiquitous in modern daily life. However, comprehensively evaluating CG quality remains challenging for two reasons: First, existing CG datasets lack systematic descriptions of rendering quality; and second existing CG quality assessment methods cannot provide reasonable text-based explanations. To address these issues, we first identify six key perceptual dimensions of CG quality from the user perspective and construct a dataset of 3500 CG images with corresponding quality descriptions. Each description covers CG style, content, and perceived quality along the selected dimensions. Furthermore, we use a subset of the dataset to build several question-answer benchmarks based on the descriptions in order to evaluate the responses of existing Vision Language Models (VLMs). We find that current VLMs are not sufficiently accurate in judging fine-grained CG quality, but that descriptions of visually similar images can significantly improve a VLM's understanding of a given CG image. Motivated by this observation, we adopt retrieval-augmented generation and propose a two-stream retrieval framework that effectively enhances the CG quality assessment capabilities of VLMs. Experiments on several representative VLMs demonstrate that our method substantially improves their performance on CG quality assessment.

图像质量评估视觉语言模型检索增强计算机图形

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。