arXiv:2608.14226cs.CV2026-08

自动发现文本生成图像模型中可编辑且多样化的语义,提升编辑效率。

RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models

论文配图:RankT2I: A Submodular Framework for Discovering Interpretable and Diverse Semantics in Text-to-Image Models
图 1 · 摘自论文原文
  • 用子模函数建模语义选择,兼顾相关性、可编辑性和多样性。
  • 在多个视觉领域上识别出更广泛有效的可编辑语义,优于现有方法。
  • 无需训练、不依赖特定模型,适用于扩散模型和FLUX类生成系统。

近年来,文本到图像(T2I)模型在图像生成与编辑领域取得突破性进展。然而,如何识别模型能成功编辑的语义仍是一大挑战。现有方法通常需用户手动指定待修改语义,过程耗时且依赖反复试错。本文提出RankT2I,一种无需训练、模型无关的自动化框架,用于在扩散模型与FLUX-based模型中发现可编辑语义。给定一个视觉领域,我们首先利用多模态视觉-语言模型获取大量候选语义。随后将语义发现建模为集合选择问题,采用子模目标函数筛选出相关、可编辑且多样化的语义。该方法显著提升了跨多个领域的文本到图像编辑语义发现效率,并优于现有方法。

原文摘要 · Abstract (English)

Recent advances in text-to-image (T2I) models have revolutionized the field of image generation and editing. However, identifying semantics that a T2I model can successfully edit in an image continues to be a challenging task. Most existing approaches require users to manually specify semantics to modify a particular image, a time-consuming process that often involves extensive trial and error. In this paper, we present RankT2I, a novel, training-free, and model-agnostic framework that automates the discovery of editable semantics in diffusion and FLUX-based models. Given a visual domain, we first utilize a multimodal vision-language model to gather a broad set of candidate semantics. We then frame semantic discovery as a set selection problem and use a submodular objective to identify semantics that are relevant, editable, and diverse. Our method helps users efficiently identify a wide range of semantics for text-to-image editing models across several domains while outperforming existing methods.

文本生成图像语义发现扩散模型可编辑性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。