用最优控制理论提升文本生成多物体图像的准确性与一致性。
FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation
- 将流匹配转化为随机最优控制,仅用一个超参数调节对象保真度与分离性。
- 提出无需训练的测试时控制器和轻量微调方法,显著改善多主体对齐。
- 适用于主流模型,可推广到未见过的提示词,适合图像生成研究者使用。
文本到图像(T2I)模型在单主体提示下表现优异,但在多主体场景中常出现属性泄露、身份纠缠和主体遗漏问题。本文提出一种理论严谨的框架,将流匹配(FM)建模为随机最优控制(SOC),实现保真度与对象中心状态分离/绑定一致性的单一超参数调控。在此框架下,推导出两种与架构无关的算法:(i) 无训练的测试时控制器,通过一次前向更新扰动基础速度;(ii) 伴随匹配(Adjoint Matching),一种轻量微调规则,回归控制网络至反向伴随信号。该框架统一了已有注意力启发式方法,通过流-扩散对应关系拓展至扩散模型,并首次提供专为多主体保真度设计的微调路径。此外,引入FOCUS(Flow Optimal Control for Unentangled Subjects),一种兼容两种算法的概率性注意力绑定目标。实验表明,在Stable Diffusion 3.5和FLUX.1上,两种算法均持续提升多主体对齐效果,同时保持基模型风格;测试时控制在消费级显卡上高效运行,微调模型可泛化至未见提示词。
原文摘要 · Abstract (English)
Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. We present a principled theoretical framework that steers sampling toward multi-subject fidelity by casting flow matching (FM) as stochastic optimal control (SOC), yielding a single hyperparameter controlled trade-off between fidelity and object-centric state separation / binding consistency. Within this framework, we derive two architecture-agnostic algorithms: (i) a training-free test-time controller that perturbs the base velocity with a single-pass update, and (ii) Adjoint Matching, a lightweight fine-tuning rule that regresses a control network to a backward adjoint signal. The same formulation unifies prior attention heuristics, extends to diffusion models via a flow--diffusion correspondence, and provides the first fine-tuning route explicitly designed for multi-subject fidelity. In addition, we also introduce FOCUS (Flow Optimal Control for Unentangled Subjects), a probabilistic attention-binding objective compatible with both algorithms. Empirically, on Stable Diffusion 3.5 and FLUX.1, both algorithms consistently improve multi-subject alignment while maintaining base-model style; test-time control runs efficiently on commodity GPUs, and fine-tuned models generalize to unseen prompts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。