用极简设计实现扩散模型通用图像控制,性能超越专用方法。
OminiControl: Minimal and Universal Control for Diffusion Transformer
- 仅增加0.1%参数,复用DiT自身编码器与注意力块
- 统一处理条件与图像标记,灵活支持多种控制任务
- 动态位置编码适配对齐与非对齐任务,适合高效生成系统
我们提出OminiControl,一种重新思考图像条件如何融入扩散Transformer(DiT)架构的新方法。现有图像条件方法或引入大量参数开销,或仅对特定控制任务有效,限制了实际应用的通用性。OminiControl通过三项关键创新解决此问题:(1) 极简架构设计,复用DiT自身VAE编码器与Transformer块,仅增加0.1%额外参数;(2) 统一序列处理策略,将条件标记与图像标记融合,实现灵活的标记交互;(3) 动态位置编码机制,可适应空间对齐与非对齐控制任务。大量实验表明,该简化方法在多项条件生成任务中不仅达到甚至超过专用方法的性能。为克服主体驱动生成中的数据局限,我们还构建了Subjects200K,一个使用DiT模型自动生成的、具有身份一致性的大规模图像对数据集。本工作证明,无需复杂架构即可实现高效图像控制,为高效且通用的图像生成系统开辟新路径。
原文摘要 · Abstract (English)
We present OminiControl, a novel approach that rethinks how image conditions are integrated into Diffusion Transformer (DiT) architectures. Current image conditioning methods either introduce substantial parameter overhead or handle only specific control tasks effectively, limiting their practical versatility. OminiControl addresses these limitations through three key innovations: (1) a minimal architectural design that leverages the DiT's own VAE encoder and transformer blocks, requiring just 0.1% additional parameters; (2) a unified sequence processing strategy that combines condition tokens with image tokens for flexible token interactions; and (3) a dynamic position encoding mechanism that adapts to both spatially-aligned and non-aligned control tasks. Our extensive experiments show that this streamlined approach not only matches but surpasses the performance of specialized methods across multiple conditioning tasks. To overcome data limitations in subject-driven generation, we also introduce Subjects200K, a large-scale dataset of identity-consistent image pairs synthesized using DiT models themselves. This work demonstrates that effective image control can be achieved without architectural complexity, opening new possibilities for efficient and versatile image generation systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。