arXiv:2510.22107cs.CVcs.AI2025-10NeurIPS被引 1

用隐式图结构让图像生成更多样,自动捕捉提示不确定性

Discovering Latent Graphs with GFlowNets for Diverse Conditional Image Generation

  • 将提示分解为多个隐式表征,通过流网络采样生成多样性图像
  • 在自然与医学图像上提升生成多样性和保真度,跨任务表现优异
  • 适合需要多解输出的生成场景,如医疗影像、创意设计

捕捉多样性在条件化和基于提示的图像生成中至关重要,尤其当条件存在不确定性时可能引发多种合理输出。传统方法依赖随机种子或提示词修改,难以区分有意义差异。本文提出Rainbow框架,适用于任意预训练条件生成模型,通过将输入条件分解为多个捕获不确定性的潜在表征,生成多样化且合理的图像。首先,在提示表示计算中引入由生成流网络(GFlowNets)参数化的潜在图;其次,利用GFlowNets强大的图采样能力,从图中生成多条轨迹,共同表征输入条件,从而产生多样化的条件表示及对应图像。在自然图像与医学图像数据集上的评估显示,Rainbow在图像合成、生成与反事实生成任务中均显著提升多样性与保真度。

原文摘要 · Abstract (English)

Capturing diversity is crucial in conditional and prompt-based image generation, particularly when conditions contain uncertainty that can lead to multiple plausible outputs. To generate diverse images reflecting this diversity, traditional methods often modify random seeds, making it difficult to discern meaningful differences between samples, or diversify the input prompt, which is limited in verbally interpretable diversity. We propose Rainbow, a novel conditional image generation framework, applicable to any pretrained conditional generative model, that addresses inherent condition/prompt uncertainty and generates diverse plausible images. Rainbow is based on a simple yet effective idea: decomposing the input condition into diverse latent representations, each capturing an aspect of the uncertainty and generating a distinct image. First, we integrate a latent graph, parameterized by Generative Flow Networks (GFlowNets), into the prompt representation computation. Second, leveraging GFlowNets' advanced graph sampling capabilities to capture uncertainty and output diverse trajectories over the graph, we produce multiple trajectories that collectively represent the input condition, leading to diverse condition representations and corresponding output images. Evaluations on natural image and medical image datasets demonstrate Rainbow's improvement in both diversity and fidelity across image synthesis, image generation, and counterfactual generation tasks.

图像生成生成模型多样性控制潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。