让难生成的概念不被忽略,通过动态调整扩散过程提升图像合成质量。
RADIANCE: Relative Adaptive Denoising with IP-Adapter for Novel Concept Enhancement

- 用CLIP反馈监控中间特征中对象与属性的出现,实时感知生成偏差。
- 通过正负提示适配器比例调节,主动修正生成轨迹中的偏移问题。
- 无需训练即可提升罕见概念生成效果,适合需要精准控制的视觉创作场景。
文本到图像(T2I)扩散模型虽取得显著进展,但在合成涉及异常属性-物体组合的稀有概念时仍易出现概念缺失或语义漂移,即主导实体掩盖目标概念。我们发现此类失败源于去噪轨迹中缺乏组合平衡,提出RADIANCE——一种无需训练的推理闭环框架。该框架在预训练模型基础上引入三个模块:(1) 组合相似性监测器(CSM),基于CLIP反馈追踪中间潜在表示中对象与属性的涌现;(2) 双向尺度控制器(BSC),利用正负IP-Adapter尺度施加“恢复力”,重新平衡偏差轨迹;(3) 反馈引导调度器(FGS),在不增加训练的前提下协调各时间步更新。进一步通过延迟适配器激活(DAA)和层间交替引导(LAG)扩展至多对象提示,防止过早概念融合。通过流水线化执行监控与去噪,RADIANCE保持竞争性延迟的同时显著提升单样本成功率与有效吞吐量。在RareBench与T2I-CompBench上的实验表明,RADIANCE在组合对齐与感知质量上持续优于现有最优基线。
原文摘要 · Abstract (English)
Text-to-image (T2I) diffusion models have achieved striking progress but still struggle to synthesize rare concepts involving unusual attribute-object pairings, often resulting in concept omission or semantic drift where a dominant entity overwhelms the generation. Tracing these failures to a lack of compositional balance during the denoising trajectory, we propose RADIANCE, a training-free framework that treats inference as a closed-loop feedback process. RADIANCE augments pretrained backbones with three modular components: (1) a Compositional Similarity Monitor (CSM) that tracks the emergence of objects and attributes in intermediate latents via CLIP-based feedback; (2) a Bidirectional Scale Controller (BSC) that applies a reactive "restoring force" using positive and negative IP-Adapter scales to rebalance biased trajectories; and (3) a Feedback Guidance Scheduler (FGS) that coordinates these updates across timesteps without additional training. We further extend the framework to multi-object prompts via Delayed Adapter Activation (DAA) and Layer-wise Alternating Guidance (LAG) to prevent premature concept fusion. By overlapping monitoring and denoising through pipelined execution, RADIANCE maintains competitive latency while significantly enhancing the per-sample success rate and effective throughput. Experiments on RareBench and T2I-CompBench demonstrate that RADIANCE consistently enhances compositional alignment and perceptual quality over state-of-the-art baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。