arXiv:2503.10697cs.CVcs.AI2025-03ICCV被引 4

用熵融合技术精准提取主体,生成无多余元素的创意图像。

Zero-Shot Subject-Centric Generation for Creative Application Using Entropy Fusion

  • 基于熵加权融合跨注意力特征,精准预测主体掩码。
  • 在纺织设计与表情包生成中,主体生成质量显著优于现有方法。
  • 结合大模型理解口语化输入,适合创意设计初学者使用。

生成模型广泛应用于视觉内容创作,但当前文本到图像模型在实际应用(如纺织图案设计、表情包生成)中常因难以分离的无关元素而受限。为此,本文提出一种可靠的主体中心图像生成框架。通过基于熵的特征加权融合方法,整合预训练文生图模型FLUX每一步采样获得的交叉注意力特征,实现精确掩码预测与主体中心生成。同时,构建基于大语言模型的智能代理框架,将用户非正式输入转化为更详细的提示词,提升图像细节表现力;并提取提示词中的核心元素,引导熵融合过程,确保生成聚焦于主要对象且无冗余成分。实验结果与用户研究验证了本方法在主体生成质量上优于现有方法或其它可能流程,证明其有效性。

原文摘要 · Abstract (English)

Generative models are widely used in visual content creation. However, current text-to-image models often face challenges in practical applications-such as textile pattern design and meme generation-due to the presence of unwanted elements that are difficult to separate with existing methods. Meanwhile, subject-reference generation has emerged as a key research trend, highlighting the need for techniques that can produce clean, high-quality subject images while effectively removing extraneous components. To address this challenge, we introduce a framework for reliable subject-centric image generation. In this work, we propose an entropy-based feature-weighted fusion method to merge the informative cross-attention features obtained from each sampling step of the pretrained text-to-image model FLUX, enabling a precise mask prediction and subject-centric generation. Additionally, we have developed an agent framework based on Large Language Models (LLMs) that translates users' casual inputs into more descriptive prompts, leading to highly detailed image generation. Simultaneously, the agents extract primary elements of prompts to guide the entropy-based feature fusion, ensuring focused primary element generation without extraneous components. Experimental results and user studies demonstrate our methods generates high-quality subject-centric images, outperform existing methods or other possible pipelines, highlighting the effectiveness of our approach.

图像生成主体提取大模型应用创意设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。