用上下文学习统一图像定制,支持多种工业场景。
IC-Custom: Diverse Image Customization via In-Context Learning
- 通过多图拼接和可学习的注册令牌实现位置感知与无位置定制融合。
- 在12K数据集上训练仅0.4%参数,人评一致性、和谐度等指标提升73%。
- 适合服装试穿、图像插入等实际工业应用,效果超越主流模型。
图像定制是工业媒体生产的关键技术,旨在生成与参考图像一致的内容。现有方法通常将定制分为位置感知与无位置两种范式,缺乏统一框架,限制了多场景应用。为此,我们提出IC-Custom,一种通过上下文学习无缝融合两类定制的统一框架。该框架将参考图像与目标图像拼接为多联画(polyptych),利用DiT的多模态注意力机制实现细粒度的标记级交互。我们提出上下文多模态注意力(ICMA)机制,采用可学习的任务导向注册令牌和边界感知位置嵌入,使模型能有效处理多样任务并区分多图配置中的输入。为弥补数据缺口,我们构建了包含8K真实世界与4K高质量合成样本的12K身份一致性数据集,避免合成数据常见的过度饱和问题。IC-Custom支持试穿、图像插入及创意知识产权定制等多种工业应用。在自建ProductBench与公开的DreamBench上的大量评估表明,其显著优于社区工作、闭源模型及前沿开源方法。在身份一致性、协调性与文本对齐等指标上,人类偏好提升约73%,且仅训练原模型0.4%的参数。
原文摘要 · Abstract (English)
Image customization, a crucial technique for industrial media production, aims to generate content that is consistent with reference images. However, current approaches conventionally separate image customization into position-aware and position-free customization paradigms and lack a universal framework for diverse customization, limiting their applications across various scenarios. To overcome these limitations, we propose IC-Custom, a unified framework that seamlessly integrates position-aware and position-free image customization through in-context learning. IC-Custom concatenates reference images with target images to a polyptych, leveraging DiT's multi-modal attention mechanism for fine-grained token-level interactions. We propose the In-context Multi-Modal Attention (ICMA) mechanism, which employs learnable task-oriented register tokens and boundary-aware positional embeddings to enable the model to effectively handle diverse tasks and distinguish between inputs in polyptych configurations. To address the data gap, we curated a 12K identity-consistent dataset with 8K real-world and 4K high-quality synthetic samples, avoiding the overly glossy, oversaturated look typical of synthetic data. IC-Custom supports various industrial applications, including try-on, image insertion, and creative IP customization. Extensive evaluations on our proposed ProductBench and the publicly available DreamBench demonstrate that IC-Custom significantly outperforms community workflows, closed-source models, and state-of-the-art open-source approaches. IC-Custom achieves about 73\% higher human preference across identity consistency, harmony, and text alignment metrics, while training only 0.4\% of the original model parameters. Project page: https://liyaowei-stu.github.io/project/IC_Custom
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。