arXiv:2412.03178cs.AIcs.CV2024-12CVPR被引 11

首次量化文本生成图像的不确定性,提升模型可信度。

Towards Understanding and Quantifying Uncertainty for Text-to-Image Generation

  • 用视觉语言模型对比生成图与提示语义差异,评估不确定性
  • 可分离随机与认知不确定性,优于传统图像空间方法
  • 适用于检测偏见、版权保护等场景,适合可信AI研究者

文本到图像生成模型中的不确定性量化对理解模型行为、提升输出可靠性至关重要。本文首次针对提示词对文本生成图像模型的不确定性进行量化与评估。在沿用现有图像空间不确定性度量方法的基础上,提出基于提示词的不确定性估计方法(PUNC),利用大视觉语言模型(LVLM)为生成图像生成描述,并将其与原始提示词在更语义化的文本空间中对比,以更准确捕捉由提示语义和生成图像引发的不确定性。PUNC通过精确率与召回率实现对随机不确定性与认知不确定性的解耦,这是传统图像空间方法无法做到的。大量实验表明,PUNC在多种设置下均超越当前最优不确定性估计技术。该方法可用于偏见检测、版权保护及分布外检测等应用。同时,我们构建了一个包含文本提示与生成结果对的综合性数据集,以推动生成模型不确定性研究。结果表明,PUNC不仅性能优异,还支持新应用场景,助力提升文本生成图像模型的可信性。

原文摘要 · Abstract (English)

Uncertainty quantification in text-to-image (T2I) generative models is crucial for understanding model behavior and improving output reliability. In this paper, we are the first to quantify and evaluate the uncertainty of T2I models with respect to the prompt. Alongside adapting existing approaches designed to measure uncertainty in the image space, we also introduce Prompt-based UNCertainty Estimation for T2I models (PUNC), a novel method leveraging Large Vision-Language Models (LVLMs) to better address uncertainties arising from the semantics of the prompt and generated images. PUNC utilizes a LVLM to caption a generated image, and then compares the caption with the original prompt in the more semantically meaningful text space. PUNC also enables the disentanglement of both aleatoric and epistemic uncertainties via precision and recall, which image-space approaches are unable to do. Extensive experiments demonstrate that PUNC outperforms state-of-the-art uncertainty estimation techniques across various settings. Uncertainty quantification in text-to-image generation models can be used on various applications including bias detection, copyright protection, and OOD detection. We also introduce a comprehensive dataset of text prompts and generation pairs to foster further research in uncertainty quantification for generative models. Our findings illustrate that PUNC not only achieves competitive performance but also enables novel applications in evaluating and improving the trustworthiness of text-to-image models.

不确定性量化文本生成图像可信AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。