arXiv:2602.17200cs.CV2026-02被引 1

通过几何采样提升文本生成图像的多样性,不牺牲质量与语义一致性。

GASS: Geometry-Aware Spherical Sampling for Disentangled Diversity Enhancement in Text-to-Image Generation

  • 基于几何分解,分离提示相关与无关的变异源,分别控制。
  • 在多个模型和数据集上实现多样性的显著提升,保持图像质量。
  • 适合需要丰富生成结果的创意设计、内容生成场景。

尽管现代文本到图像(T2I)生成模型具有较高的语义对齐能力,但仍难以从同一提示中生成多样化图像。本文从几何视角增强生成多样性,提出几何感知球面采样(GASS)。不同于依赖熵引导的方法,GASS通过分解CLIP嵌入中的多样性度量,将变化分为两个正交方向:文本嵌入方向(捕获与提示相关的语义变化),以及一个确定的正交方向(捕获提示无关的变化,如背景)。基于此分解,GASS在两个轴上扩大生成图像嵌入的几何投影分布,并通过沿生成轨迹扩展预测来指导采样过程。在多个冻结的T2I骨干网络(包括U-Net和DiT,扩散与流模型)及基准测试上的实验表明,该方法有效提升了解耦多样性,且对图像保真度和语义对齐影响极小。

原文摘要 · Abstract (English)

Despite high semantic alignment, modern text-to-image (T2I) generative models still struggle to synthesize diverse images from a given prompt. In this work, we enhance the T2I diversity through a geometric lens. Unlike most existing methods that rely primarily on entropy-based guidance to increase sample dissimilarity, we introduce Geometry-Aware Spherical Sampling (GASS) to enhance diversity by explicitly controlling both prompt-dependent and prompt-independent sources of variation. Specifically, we decompose the diversity measure in CLIP embeddings using two orthogonal directions: the text embedding, which captures semantic variation related to the prompt, and an identified orthogonal direction that captures prompt-independent variation (e.g., backgrounds). Based on this decomposition, GASS increases the geometric projection spread of generated image embeddings along both axes and guides the T2I sampling process via expanded predictions along the generation trajectory. Our experiments on different frozen T2I backbones (U-Net and DiT, diffusion and flow) and benchmarks demonstrate the effectiveness of disentangled diversity enhancement with minimal impact on image fidelity and semantic alignment.

图像生成多样性增强几何采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。