TIDE让图像生成既保主体又听指令,无需微调
TIDE: Achieving Balanced Subject-Driven Image Generation via Target-Instructed Diffusion Enhancement
- 用参考图+指令+目标图三元组建模主体变化动态
- 通过量化对比生成优/劣样本,自动学习平衡策略
- 适合需要精准控制主体形象的文本生成任务
主体驱动图像生成(SDIG)旨在根据文本指令操控图像中的特定主体,是推动文生图扩散模型发展的关键任务。该任务需调和主体身份保持与指令动态适配之间的矛盾,现有方法未能有效解决。本文提出目标指导扩散增强(TIDE)框架,通过目标监督与偏好学习,在无需测试时微调的情况下化解这一矛盾。TIDE首创目标监督三元组对齐,利用(参考图像、指令、目标图像)三元组建模主体适应过程。该方法采用直接主体扩散(DSD)目标,通过定量指标系统生成并评估“胜出”(平衡保真与合规)和“失败”(失真)样本进行训练,实现对最优保真-合规平衡的隐式奖励建模。在标准基准上的实验表明,TIDE在生成主体忠实输出的同时保持指令合规性,优于多个基线方法。其通用性还体现在结构条件生成、图到图生成及文本-图像插值等多样化任务中。代码已开源。
原文摘要 · Abstract (English)
Subject-driven image generation (SDIG) aims to manipulate specific subjects within images while adhering to textual instructions, a task crucial for advancing text-to-image diffusion models. SDIG requires reconciling the tension between maintaining subject identity and complying with dynamic edit instructions, a challenge inadequately addressed by existing methods. In this paper, we introduce the Target-Instructed Diffusion Enhancing (TIDE) framework, which resolves this tension through target supervision and preference learning without test-time fine-tuning. TIDE pioneers target-supervised triplet alignment, modelling subject adaptation dynamics using a (reference image, instruction, target images) triplet. This approach leverages the Direct Subject Diffusion (DSD) objective, training the model with paired "winning" (balanced preservation-compliance) and "losing" (distorted) targets, systematically generated and evaluated via quantitative metrics. This enables implicit reward modelling for optimal preservation-compliance balance. Experimental results on standard benchmarks demonstrate TIDE's superior performance in generating subject-faithful outputs while maintaining instruction compliance, outperforming baseline methods across multiple quantitative metrics. TIDE's versatility is further evidenced by its successful application to diverse tasks, including structural-conditioned generation, image-to-image generation, and text-image interpolation. Our code is available at https://github.com/KomJay520/TIDE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。