统一框架实现多类别图像插入,保持参考图像精度与生成质量。
InsertFuse: A Unified Framework for Multi-Category Reference-Guided Image Insertion

- 分阶段训练专家模型,再通过策略蒸馏融合能力
- 在任意插入基准上达到当前最优性能,参考保真度高
- 适合需要多类型图像插入的生成系统开发者
我们提出 InsertFuse,一种面向多类别参考引导图像插入的统一框架。其核心思想是将类别专属知识学习与跨类别能力整合相分离。首先为不同插入类别训练专用专家模型,随后引入插入策略蒸馏(IOPD),将各专家能力凝聚到单一学生模型中。通过在学生访问的状态下查询匹配专家,IOPD保留类别特异性插入行为,同时缓解直接联合训练带来的跨类别干扰。为增强空间控制,提出基于令牌对齐的几何条件(TAGC),将掩码导出的几何线索映射至视觉令牌网格;并引入区域平衡流匹配,分别归一化插入区域内外的预测误差,避免背景主导和尺度依赖的监督。进一步提出参考CFG,在固定场景与几何条件下隔离并强化视觉参考的引导作用,再由IOPD将此增强引导转移至统一学生模型。在公开的 AnyInsertion 基准和自建多类别测试集上的大量实验表明,该方法在多数指标上达到当前最优,展现出多样插入类别下的强参考保真度与高质量生成效果。
原文摘要 · Abstract (English)
We present InsertFuse, a unified framework for multi-category reference-guided image insertion. Its key idea is to decouple category-specific expertise learning from cross-category capability consolidation. InsertFuse first trains specialized experts for different insertion categories and then introduces Insertion On-Policy Distillation (IOPD) to consolidate their capabilities into a single student. By querying the matched expert at states visited by the student, IOPD preserves category-specific insertion behavior while mitigating the cross-category interference caused by direct joint training. To improve spatial control, we propose Token-Aligned Geometry Conditioning (TAGC), which maps mask-derived geometric cues to the visual token grid, and Region-Balanced Flow Matching, which separately normalizes prediction errors inside and outside the insertion region to prevent background-dominated and scale-dependent supervision. We further introduce Reference CFG to isolate and strengthen the guidance induced by the visual reference under fixed scene and geometry conditions, with IOPD transferring this enhanced supervision into the unified student. Extensive experiments on the public AnyInsertion benchmark and our multi-category test set demonstrate state-of-the-art performance on most metrics, showing strong reference fidelity and generation quality across diverse insertion categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。