arXiv:2504.08714cs.CVcs.CL2025-04EMNLP

用大模型提升图像中物体互动细节生成质量

Generating Fine Details of Entity Interactions

  • 将互动分解为细粒度概念,通过大模型迭代优化
  • 在1000个精细提示上实现显著图像质量提升
  • 适合需要复杂物体交互生成的研究者与开发者

当前文本到图像模型虽能生成高质量物体为中心的图像,但在表现物体间丰富互动方面仍显不足,主要因稀有互动数据和基准匮乏。本文提出一种新方法,利用多模态大语言模型(MLLM)构建并改进互动丰富的图像生成。我们构建了包含1000个由大语言模型生成的细粒度图像生成提示的数据集 exttt{data},覆盖功能/动作型互动、多主体互动及组合空间关系。针对生成挑战,提出分解-增强精炼流程 exttt{model}:先由大模型拆解互动为细粒度概念,再用MLLM评估生成图像,最后通过部分扩散去噪进行针对性修正。自动与人工评估均显示图像质量显著提升,证明了增强推理策略的有效性。

原文摘要 · Abstract (English)

Recent text-to-image models excel at generating high-quality object-centric images from instructions. However, images should also encapsulate rich interactions between objects, where existing models often fall short, likely due to limited training data and benchmarks for rare interactions. This paper explores a novel application of Multimodal Large Language Models (MLLMs) to benchmark and enhance the generation of interaction-rich images. We introduce \data, an interaction-focused dataset with 1000 LLM-generated fine-grained prompts for image generation covering (1) functional and action-based interactions, (2) multi-subject interactions, and (3) compositional spatial relationships. To address interaction-rich generation challenges, we propose a decomposition-augmented refinement procedure. Our approach, \model, leverages LLMs to decompose interactions into finer-grained concepts, uses an MLLM to critique generated images, and applies targeted refinements with a partial diffusion denoising process. Automatic and human evaluations show significantly improved image quality, demonstrating the potential of enhanced inference strategies.

图像生成多模态大模型互动细节

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。