arXiv:2508.09575cs.CV2025-08ICCV

无需训练即可精准控制图像生成中的姿态与外观细节。

Dual Recursive Feedback on Generation and Appearance Latents for Pose-Robust Text-to-Image Diffusion

  • 通过外观与生成双反馈循环优化中间潜在表示。
  • 在不改变结构的前提下,实现跨类姿态迁移如虎身换人动。
  • 适合需要精细控制图像生成的视觉创作与设计场景。

可控文本到图像(T2I)扩散模型如Ctrl-X和FreeControl已在无需额外模块训练的情况下展现出良好的空间与外观控制能力。然而,这些模型常难以准确保留空间结构,且无法捕捉物体姿态与场景布局等细粒度条件。为此,我们提出一种无需训练的双重递归反馈(DRF)系统,能有效将控制条件反映在生成过程中。该系统包含外观反馈与生成反馈,通过递归精修中间潜在表示,使潜在特征更准确地体现给定的外观信息与用户意图。这一双更新机制引导潜在表示向可靠流形收敛,实现结构与外观属性的有效融合。方法可在类间不变的结构-外观融合场景下实现精细生成,例如将人体动作迁移到老虎形态上。大量实验表明,该方法能生成高质量、语义连贯且结构一致的图像。代码已开源:https://github.com/jwonkm/DRF。

原文摘要 · Abstract (English)

Recent advancements in controllable text-to-image (T2I) diffusion models, such as Ctrl-X and FreeControl, have demonstrated robust spatial and appearance control without requiring auxiliary module training. However, these models often struggle to accurately preserve spatial structures and fail to capture fine-grained conditions related to object poses and scene layouts. To address these challenges, we propose a training-free Dual Recursive Feedback (DRF) system that properly reflects control conditions in controllable T2I models. The proposed DRF consists of appearance feedback and generation feedback that recursively refines the intermediate latents to better reflect the given appearance information and the user's intent. This dual-update mechanism guides latent representations toward reliable manifolds, effectively integrating structural and appearance attributes. Our approach enables fine-grained generation even between class-invariant structure-appearance fusion, such as transferring human motion onto a tiger's form. Extensive experiments demonstrate the efficacy of our method in producing high-quality, semantically coherent, and structurally consistent image generations. Our source code is available at https://github.com/jwonkm/DRF.

扩散模型姿态控制图像生成零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。