arXiv:2508.17760cs.CVcs.CL2025-08

通过双控机制提升文本生成图像中实体与互动的准确性。

CEIDM: A Controlled Entity and Interaction Diffusion Model for Enhanced Text-to-Image Generation

  • 用大模型挖掘文本隐含的实体互动关系,引导生成更合理图像。
  • 对动作特征聚类并偏移,增强动作语义理解与细节还原。
  • 设计实体控制网络,精准生成实体掩码并提升图像质量。

在文本到图像(T2I)生成中,实体的复杂性及其相互作用给基于扩散模型的方法带来挑战:如何有效控制实体及其交互以生成高质量图像。为此,我们提出CEIDM,一种基于扩散模型的双控图像生成方法,分别控制实体与交互。首先,利用大语言模型(LLMs)的思维链方法挖掘实体间的隐含互动关系,提取丰富合理的交互逻辑,指导扩散模型生成更符合现实逻辑、互动更合理的图像。其次,提出交互动作聚类与偏移方法,对文本提示中的动作特征进行聚类和偏移处理,构建全局与局部双向偏移,增强动作语义理解与细节补全,使模型对交互‘动作’的概念理解更准确,生成图像中动作更精确。最后,设计实体控制网络,通过语义引导生成实体掩码,结合多尺度卷积网络增强实体特征,并采用动态网络融合特征,有效控制实体并显著提升图像质量。实验表明,所提方法在实体控制与交互控制方面均优于现有代表性方法。

原文摘要 · Abstract (English)

In Text-to-Image (T2I) generation, the complexity of entities and their intricate interactions pose a significant challenge for T2I method based on diffusion model: how to effectively control entity and their interactions to produce high-quality images. To address this, we propose CEIDM, a image generation method based on diffusion model with dual controls for entity and interaction. First, we propose an entity interactive relationships mining approach based on Large Language Models (LLMs), extracting reasonable and rich implicit interactive relationships through chain of thought to guide diffusion models to generate high-quality images that are closer to realistic logic and have more reasonable interactive relationships. Furthermore, We propose an interactive action clustering and offset method to cluster and offset the interactive action features contained in each text prompts. By constructing global and local bidirectional offsets, we enhance semantic understanding and detail supplementation of original actions, making the model's understanding of the concept of interactive "actions" more accurate and generating images with more accurate interactive actions. Finally, we design an entity control network which generates masks with entity semantic guidance, then leveraging multi-scale convolutional network to enhance entity feature and dynamic network to fuse feature. It effectively controls entities and significantly improves image quality. Experiments show that the proposed CEIDM method is better than the most representative existing methods in both entity control and their interaction control.

文本生成图像扩散模型实体控制交互生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。