arXiv:2505.20053cs.CVcs.AI2025-05被引 7

用多模态大模型实时修正图像生成中的语义错误。

Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion

  • 引入多模态大模型作为生成过程中的语义观察者
  • 在少数扩散步骤中实现显著的语义对齐提升
  • 适合需要精准控制生成内容的场景

扩散模型已成为文本到图像生成的主流架构,在视觉质量和提示可控性方面取得了显著进展。然而,当前的推理流程普遍缺乏可解释的语义监督与修正机制,多数方法仅依赖最终图像的后处理评分、提示过滤或启发式重采样策略,无法为生成轨迹提供有效指导。因此,模型常出现物体混淆、空间错误、计数不准和语义缺失等问题,严重损害提示-图像对齐度与图像质量。为此,我们提出一种新框架——MLLM引导的语义修正乒乓提前扩散(PPAD),首次在推理阶段引入多模态大语言模型(MLLM)作为语义观察者。该框架对中间生成结果进行实时分析,识别潜在语义不一致,并将反馈转化为可调控信号,主动引导剩余去噪步骤。该方法支持仅推理与训练增强两种设置,且仅需极少扩散步骤即可完成语义修正,具备强泛化性与可扩展性。大量实验表明,PPAD在多个指标上均有显著提升。

原文摘要 · Abstract (English)

Diffusion models have become the mainstream architecture for text-to-image generation, achieving remarkable progress in visual quality and prompt controllability. However, current inference pipelines generally lack interpretable semantic supervision and correction mechanisms throughout the denoising process. Most existing approaches rely solely on post-hoc scoring of the final image, prompt filtering, or heuristic resampling strategies-making them ineffective in providing actionable guidance for correcting the generative trajectory. As a result, models often suffer from object confusion, spatial errors, inaccurate counts, and missing semantic elements, severely compromising prompt-image alignment and image quality. To tackle these challenges, we propose MLLM Semantic-Corrected Ping-Pong-Ahead Diffusion (PPAD), a novel framework that, for the first time, introduces a Multimodal Large Language Model (MLLM) as a semantic observer during inference. PPAD performs real-time analysis on intermediate generations, identifies latent semantic inconsistencies, and translates feedback into controllable signals that actively guide the remaining denoising steps. The framework supports both inference-only and training-enhanced settings, and performs semantic correction at only extremely few diffusion steps, offering strong generality and scalability. Extensive experiments demonstrate PPAD's significant improvements.

文本生成图像扩散模型多模态大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。