无需训练即可实时对齐用户偏好,实现精准可控的文生图生成。
Instant Preference Alignment for Text-to-Image Diffusion Models
- 利用多模态大模型解析参考图与提示词,自动提取全局偏好信号。
- 通过关键词控制与局部注意力调制,实现无训练图像生成引导。
- 支持多轮交互,适合对话式创作与个性化图像生成场景。
文本到图像(T2I)生成极大拓展了创作表达能力,但实现实时、无需训练的偏好对齐仍具挑战。现有方法常依赖静态预收集偏好或微调,难以适应动态变化的用户意图。本文提出一种基于多模态大语言模型(MLLM)先验的无训练框架,将任务解耦为偏好理解与偏好引导生成两部分。在偏好理解阶段,利用MLLM从参考图像中自动提取全局偏好信号,并通过结构化指令设计增强原始提示词,覆盖更广且更精细的用户偏好。在偏好引导生成阶段,融合全局关键词控制与局部区域感知的交叉注意力调制,无需额外训练即可精确引导扩散模型,在整体属性与局部细节上实现对齐。整个框架支持多轮交互式精炼,实现实时、上下文感知的图像生成。在Viper数据集及自建基准上的实验表明,该方法在定量指标与人工评估中均优于现有方法,为对话式生成与MLLM-扩散模型融合开辟新路径。
原文摘要 · Abstract (English)
Text-to-image (T2I) generation has greatly enhanced creative expression, yet achieving preference-aligned generation in a real-time and training-free manner remains challenging. Previous methods often rely on static, pre-collected preferences or fine-tuning, limiting adaptability to evolving and nuanced user intents. In this paper, we highlight the need for instant preference-aligned T2I generation and propose a training-free framework grounded in multimodal large language model (MLLM) priors. Our framework decouples the task into two components: preference understanding and preference-guided generation. For preference understanding, we leverage MLLMs to automatically extract global preference signals from a reference image and enrich a given prompt using structured instruction design. Our approach supports broader and more fine-grained coverage of user preferences than existing methods. For preference-guided generation, we integrate global keyword-based control and local region-aware cross-attention modulation to steer the diffusion model without additional training, enabling precise alignment across both global attributes and local elements. The entire framework supports multi-round interactive refinement, facilitating real-time and context-aware image generation. Extensive experiments on the Viper dataset and our collected benchmark demonstrate that our method outperforms prior approaches in both quantitative metrics and human evaluations, and opens up new possibilities for dialog-based generation and MLLM-diffusion integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。