分析文本到图像模型的提示词与生成结果关联,发现潜在偏差。
Is What You Ask For What You Get? Investigating Concept Associations in Text-to-Image Models
- 用可解释概念刻画图文模型的条件分布
- 发现用户自定义与真实世界提示存在显著差异
- 开源可视化工具,适合非技术用户审计模型
文本到图像(T2I)模型正广泛应用于现实场景,亟需对其生成内容进行审计以确保符合任务需求。然而,如何以人类可理解的方式系统检查提示词与生成结果之间的关联仍具挑战。为此,我们提出 Concept2Concept 框架,通过可解释概念和基于这些概念定义的度量,表征视觉语言模型的条件分布。该表征使我们能够审计模型及提示-数据集。通过多个案例研究,我们分析了用户自定义分布与真实世界经验分布的条件关系。最后,我们将 Concept2Concept 实现为开源交互式可视化工具,便于非技术用户使用。演示地址:https://tinyurl.com/Concept2ConceptDemo。
原文摘要 · Abstract (English)
Text-to-image (T2I) models are increasingly used in impactful real-life applications. As such, there is a growing need to audit these models to ensure that they generate desirable, task-appropriate images. However, systematically inspecting the associations between prompts and generated content in a human-understandable way remains challenging. To address this, we propose Concept2Concept, a framework where we characterize conditional distributions of vision language models using interpretable concepts and metrics that can be defined in terms of these concepts. This characterization allows us to use our framework to audit models and prompt-datasets. To demonstrate, we investigate several case studies of conditional distributions of prompts, such as user-defined distributions or empirical, real-world distributions. Lastly, we implement Concept2Concept as an open-source interactive visualization tool to facilitate use by non-technical end-users. A demo is available at https://tinyurl.com/Concept2ConceptDemo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。