通过噪声微调生成高度相关且多样的图像,仅需10分钟训练。
NOFT: Test-Time Noise Finetune via Information Bottleneck for Highly Correlated Asset Creation
- 在测试时用信息瓶颈优化噪声,实现可控图像变化。
- 仅用14K参数和10分钟训练,即可生成高保真图像变体。
- 适合需要快速生成2D/3D AIGC资产的设计师与开发者。
扩散模型为文本到图像(T2I)和图像到图像(I2I)生成提供了强大工具。近期研究聚焦于拓扑与纹理控制,如ControlNet、IP-Adapter、Ctrl-X和DSG,这些方法依赖外部信号或扩散特征操作实现高保真可控编辑。对于多样性,通常直接选择不同噪声隐变量。然而,扩散噪声能隐式表征对应图像的拓扑与纹理流形,也是权衡内容保持与可控变化的有效平台。现有T2I与I2I扩散方法未探索压缩上下文隐变量中的信息。本文首次提出一种即插即用的噪声微调模块NOFT,集成至Stable Diffusion中,用于生成高度相关且多样化的图像。通过最优传输信息瓶颈(OT-IB)对种子噪声或反向噪声进行微调,仅需约14K可训练参数和10分钟训练时间。实验表明,测试时的NOFT能有效生成兼顾拓扑与纹理对齐的高保真图像变体。全面实验证明,NOFT是一种高效通用的重想象方法,适用于文本或图像引导下的2D/3D AIGC资产快速微调。
原文摘要 · Abstract (English)
The diffusion model has provided a strong tool for implementing text-to-image (T2I) and image-to-image (I2I) generation. Recently, topology and texture control are popular explorations, e.g., ControlNet, IP-Adapter, Ctrl-X, and DSG. These methods explicitly consider high-fidelity controllable editing based on external signals or diffusion feature manipulations. As for diversity, they directly choose different noise latents. However, the diffused noise is capable of implicitly representing the topological and textural manifold of the corresponding image. Moreover, it's an effective workbench to conduct the trade-off between content preservation and controllable variations. Previous T2I and I2I diffusion works do not explore the information within the compressed contextual latent. In this paper, we first propose a plug-and-play noise finetune NOFT module employed by Stable Diffusion to generate highly correlated and diverse images. We fine-tune seed noise or inverse noise through an optimal-transported (OT) information bottleneck (IB) with around only 14K trainable parameters and 10 minutes of training. Our test-time NOFT is good at producing high-fidelity image variations considering topology and texture alignments. Comprehensive experiments demonstrate that NOFT is a powerful general reimagine approach to efficiently fine-tune the 2D/3D AIGC assets with text or image guidance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。