统一控制图像主体、风格与结构,实现更灵活的生成。
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
- 用自适应记忆模块分离并存储不同条件先验
- 在多个基准上生成质量优于现有方法
- 适合需要多维度控制的图像创作场景
当前图像生成方法常孤立处理主体、风格和结构控制,导致特征混淆且任务迁移能力差。本文提出3SGen,一个统一框架,在单模型中完成三类条件生成。该框架采用配备可学习语义查询的多模态大模型对齐图文语义,并通过变分自编码器分支保留细节。核心是自适应任务特定记忆(ATM)模块,通过轻量级门控机制与可扩展记忆项,动态解耦、存储并检索身份、纹理、空间布局等条件先验。该设计有效减少任务间干扰,天然支持组合输入。此外,我们构建了3SGen-Bench,一个标准化评估基准,用于衡量跨任务保真度与可控性。在所提基准及其他公开数据集上的大量实验表明,3SGen在多种图像生成任务中均表现优异。
原文摘要 · Abstract (English)
Recent image generation approaches often address subject, style, and structure-driven conditioning in isolation, leading to feature entanglement and limited task transferability. In this paper, we introduce 3SGen, a task-aware unified framework that performs all three conditioning modes within a single model. 3SGen employs an MLLM equipped with learnable semantic queries to align text-image semantics, complemented by a VAE branch that preserves fine-grained visual details. At its core, an Adaptive Task-specific Memory (ATM) module dynamically disentangles, stores, and retrieves condition-specific priors, such as identity for subjects, textures for styles, and spatial layouts for structures, via a lightweight gating mechanism along with several scalable memory items. This design mitigates inter-task interference and naturally scales to compositional inputs. In addition, we propose 3SGen-Bench, a unified image-driven generation benchmark with standardized metrics for evaluating cross-task fidelity and controllability. Extensive experiments on our proposed 3SGen-Bench and other public benchmarks demonstrate our superior performance across diverse image-driven generation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。