arXiv:2506.20832cs.CVcs.AI2025-06中稿 · IEEE Transactions …被引 3

用视觉语言模型从扩散模型生成的超分辨率图像中挑选最可信的样本。

Leveraging Vision-Language Models to Select Trustworthy Super-Resolution Samples Generated by Diffusion Models

  • 利用VLM对扩散模型生成的图像进行语义、质量与伪影评估。
  • 提出可信度评分TWS,与人类偏好高度一致。
  • 适合关注生成结果可靠性与语义准确性的研究者。

超分辨率(SR)是一个病态逆问题,存在多种与低分辨率图像一致的可行解。传统回归模型在保真度与感知质量间权衡,常引入伪影,影响关键信息识别;而扩散模型虽能生成多样解,但如何从中选出最可信样本仍是挑战。本文提出一种基于视觉语言模型(VLMs,如BLIP-2、GPT-4o)的自动化框架,通过结构化提示评估语义正确性、视觉质量与伪影存在性,筛选出高可信度候选并集成输出。为此,提出新型可信度评分(TWS),融合三重指标:基于CLIP嵌入的语义相似性、基于边缘图的SSIM结构完整性、多级小波分解的伪影敏感性。实验表明,TWS与人类偏好强相关,在模糊与自然图像上均表现优异,且显著优于PSNR、LPIPS等传统指标。该方法为扩散超分辨率空间中的不确定性提供了可扩展、可泛化的可信度判断方案。

原文摘要 · Abstract (English)

Super-resolution (SR) is an ill-posed inverse problem with many feasible solutions consistent with a given low-resolution image. On one hand, regressive SR models aim to balance fidelity and perceptual quality to yield a single solution, but this trade-off often introduces artifacts that create ambiguity in information-critical applications such as recognizing digits or letters. On the other hand, diffusion models generate a diverse set of SR images, but selecting the most trustworthy solution from this set remains a challenge. This paper introduces a robust, automated framework for identifying the most trustworthy SR sample from a diffusion-generated set by leveraging the semantic reasoning capabilities of vision-language models (VLMs). Specifically, VLMs such as BLIP-2, GPT-4o, and their variants are prompted with structured queries to assess semantic correctness, visual quality, and artifact presence. The top-ranked SR candidates are then ensembled to yield a single trustworthy output in a cost-effective manner. To rigorously assess the validity of VLM-selected samples, we propose a novel Trustworthiness Score (TWS) a hybrid metric that quantifies SR reliability based on three complementary components: semantic similarity via CLIP embeddings, structural integrity using SSIM on edge maps, and artifact sensitivity through multi-level wavelet decomposition. We empirically show that TWS correlates strongly with human preference in both ambiguous and natural images, and that VLM-guided selections consistently yield high TWS values. Compared to conventional metrics like PSNR, LPIPS, which fail to reflect information fidelity, our approach offers a principled, scalable, and generalizable solution for navigating the uncertainty of the diffusion SR space. By aligning outputs with human expectations and semantic correctness, this work sets a new benchmark for trustworthiness in generative SR.

超分辨率扩散模型可信度评估视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。