arXiv:2506.23275cs.CVcs.AI2025-06被引 1

提出图像集生成新任务,让AI按指令生成多样且一致的多图集合。

Why Settle for One? Text-to-ImageSet Generation and Evaluation

  • 设计新任务T2IS,支持多种一致性要求的图像集生成
  • 构建涵盖596条指令的T2IS-Bench基准与评估框架
  • 无需训练的AutoT2IS方法显著优于现有通用与专用模型

尽管文本到图像模型取得显著进展,许多实际应用仍需生成具有不同一致性要求的连贯图像集。现有方法通常局限于特定领域或特定一致性方面,严重限制了其泛化能力。本文提出更具挑战性的文本到图像集(T2IS)生成问题,旨在根据用户指令生成满足多种一致性要求的图像集合。为系统研究该问题,我们首先构建T2IS-Bench,包含596条跨26个子类别的多样化指令,全面覆盖T2IS生成需求。在此基础上,提出T2IS-Eval评估框架,将用户指令转化为多维度评估标准,并通过有效评估器自适应判断生成集合与标准的一致性。随后,提出AutoT2IS——一种无需训练的框架,充分调动预训练扩散变换器的上下文能力,协调视觉元素以同时满足图像级提示对齐与集合级视觉一致性。在T2IS-Bench上的大量实验表明,各种一致性挑战均超越现有方法,而我们的AutoT2IS显著优于当前通用及专用方案。该方法还展现出推动众多未被探索的实际应用场景的潜力,证实其重要实用价值。

原文摘要 · Abstract (English)

Despite remarkable progress in Text-to-Image models, many real-world applications require generating coherent image sets with diverse consistency requirements. Existing consistent methods often focus on a specific domain with specific aspects of consistency, which significantly constrains their generalizability to broader applications. In this paper, we propose a more challenging problem, Text-to-ImageSet (T2IS) generation, which aims to generate sets of images that meet various consistency requirements based on user instructions. To systematically study this problem, we first introduce $\textbf{T2IS-Bench}$ with 596 diverse instructions across 26 subcategories, providing comprehensive coverage for T2IS generation. Building on this, we propose $\textbf{T2IS-Eval}$, an evaluation framework that transforms user instructions into multifaceted assessment criteria and employs effective evaluators to adaptively assess consistency fulfillment between criteria and generated sets. Subsequently, we propose $\textbf{AutoT2IS}$, a training-free framework that maximally leverages pretrained Diffusion Transformers' in-context capabilities to harmonize visual elements to satisfy both image-level prompt alignment and set-level visual consistency. Extensive experiments on T2IS-Bench reveal that diverse consistency challenges all existing methods, while our AutoT2IS significantly outperforms current generalized and even specialized approaches. Our method also demonstrates the ability to enable numerous underexplored real-world applications, confirming its substantial practical value. Visit our project in https://chengyou-jia.github.io/T2IS-Home.

图像生成多图一致文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。