让扩散模型同时理解多种视觉条件并灵活组合
SIGMA: Selective-Interleaved Generation with Multi-Attribute Tokens
- 用多属性标记实现文本图像序列的交错条件解析
- 在70万组交错样本上微调,支持组合编辑与属性选择性迁移
- 相比Bagel在组合任务中显著提升可控性与生成质量
近期统一模型如Bagel表明,成对图像编辑数据可有效在单一扩散Transformer中对齐多个视觉任务。然而,这些模型仍局限于单条件输入,缺乏从多种异构源合成结果的灵活性。本文提出SIGMA(Selective-Interleaved Generation with Multi-Attribute Tokens),一种统一的后训练框架,可在扩散Transformer中实现交错多条件生成。SIGMA引入选择性多属性标记,包括风格、内容、主体和身份标记,使模型能够解析并组合多个视觉条件于交错的文本-图像序列中。通过在70万组交错示例上对Bagel统一主干进行后训练,SIGMA支持组合编辑、选择性属性迁移和细粒度多模态对齐。大量实验表明,SIGMA在多样化编辑与生成任务中提升了可控性、跨条件一致性和视觉质量,在组合任务上相较Bagel有显著提升。
原文摘要 · Abstract (English)
Recent unified models such as Bagel demonstrate that paired image-edit data can effectively align multiple visual tasks within a single diffusion transformer. However, these models remain limited to single-condition inputs and lack the flexibility needed to synthesize results from multiple heterogeneous sources. We present SIGMA (Selective-Interleaved Generation with Multi-Attribute Tokens), a unified post-training framework that enables interleaved multi-condition generation within diffusion transformers. SIGMA introduces selective multi-attribute tokens, including style, content, subject, and identity tokens, which allow the model to interpret and compose multiple visual conditions in an interleaved text-image sequence. Through post-training on the Bagel unified backbone with 700K interleaved examples, SIGMA supports compositional editing, selective attribute transfer, and fine-grained multimodal alignment. Extensive experiments show that SIGMA improves controllability, cross-condition consistency, and visual quality across diverse editing and generation tasks, with substantial gains over Bagel on compositional tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。