arXiv:2409.18214cs.LG2024-09综述被引 8

系统梳理文本生成图像模型的可信性问题与研究进展

Trustworthy Text-to-Image Diffusion Models: A Timely and Focused Survey

  • 按可信属性、评估方法、基准和应用构建分类体系
  • 涵盖鲁棒性、公平性、隐私等六大非功能特性分析
  • 适合关注AI伦理与安全的研究者和开发者

文本到图像(T2I)扩散模型在图像生成方面取得显著进展,但其广泛应用也引发对可信性(如鲁棒性、公平性、安全性、隐私性、事实性和可解释性)的伦理与社会担忧,类似传统深度学习任务。由于T2I模型具有多模态特性,传统深度学习可信性研究方法难以适用。近期研究尝试通过伪造、增强、验证与评估等手段探索T2I模型的可信性,但相关深入分析仍显不足。本文提供及时且聚焦的综述,从属性、方法、基准和应用四个维度构建简洁结构化分类体系。首先介绍T2I模型基础背景,总结特定于T2I任务的关键定义与度量指标,并基于这些指标分析近期提出的方法。同时,综述现有基准与领域应用。最后,指出当前研究空白,讨论现有方法局限,并提出未来研究方向,以推动可信T2I模型的发展。我们持续更新本领域最新进展,并维护GitHub仓库:https://github.com/wellzline/Trustworthy_T2I_DMs

原文摘要 · Abstract (English)

Text-to-Image (T2I) Diffusion Models (DMs) have garnered widespread attention for their impressive advancements in image generation. However, their growing popularity has raised ethical and social concerns related to key non-functional properties of trustworthiness, such as robustness, fairness, security, privacy, factuality, and explainability, similar to those in traditional deep learning (DL) tasks. Conventional approaches for studying trustworthiness in DL tasks often fall short due to the unique characteristics of T2I DMs, e.g., the multi-modal nature. Given the challenge, recent efforts have been made to develop new methods for investigating trustworthiness in T2I DMs via various means, including falsification, enhancement, verification \& validation and assessment. However, there is a notable lack of in-depth analysis concerning those non-functional properties and means. In this survey, we provide a timely and focused review of the literature on trustworthy T2I DMs, covering a concise-structured taxonomy from the perspectives of property, means, benchmarks and applications. Our review begins with an introduction to essential preliminaries of T2I DMs, and then we summarise key definitions/metrics specific to T2I tasks and analyses the means proposed in recent literature based on these definitions/metrics. Additionally, we review benchmarks and domain applications of T2I DMs. Finally, we highlight the gaps in current research, discuss the limitations of existing methods, and propose future research directions to advance the development of trustworthy T2I DMs. Furthermore, we keep up-to-date updates in this field to track the latest developments and maintain our GitHub repository at: https://github.com/wellzline/Trustworthy_T2I_DMs

文本生成扩散模型可信AI伦理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。